Skip to content
All tags

#neural-architecture-search

4 posts

MIT 6.5940 L22–L23 Course Summary and Quantum Machine Learning: Pruning, NAS, and On-Device Training on Quantum Circuits

The last two lectures of MIT 6.5940 Fall 2024 come in two halves. The first half of Lecture 22 is a 13-page Course-Summary.pdf that redraws the course as three blocks (inference, training, application-specific) on System and Algorithm axes, then lays out the 7-item final project rubric. The second half, Quantum ML Part I, has a recording but no slides. Lecture 23 (Hanrui Wang, 99 slides) covers parameterized quantum circuits (PQCs): data encoding, parameter-shift gradients, probabilistic gradient pruning under noise (QOC), the TorchQuantum library, and QuantumNAS, which searches with a SuperCircuit and then prunes gates. It reads like a replay of the course's supernet and magnitude pruning on quantum circuits. Fall 2026 has replaced both lectures with a guest lecture.

MIT 6.5940 Lab 3: Finding a Microcontroller Model with a Supernet, Predictors, and Evolutionary Search

Lab 3 of MIT 6.5940 (Fall 2024) hands you an OFA-trained MCUNetV2 super network (more than 10^19 subnets) and the Visual Wake Words dataset. Ten questions, 100 points plus 10 bonus: implement a MACs/peak-memory efficiency predictor and a three-layer MLP accuracy predictor, write random search and evolutionary search, then find a subnet that reaches at least 92.5% accuracy under 250KB and 60M MACs. This guide maps the question structure and what each question trains. It does not include solutions.

MIT 6.5940 L8 NAS II: Scoring Architectures Without Training Them, and Putting Hardware in the Loop

Lecture 8 of MIT 6.5940 (Fall 2024) attacks the most expensive step in NAS: evaluating candidates. Training 12,800 architectures from scratch cost 22,400 GPU-hours, so the lecture walks through inherited weights, hypernetworks, ProxylessNAS's single-path training, latency lookup tables and predictors, Once-for-All's one training run for 10^19 subnets, training-free zero-shot NAS, and NAAS, which searches the network and the accelerator together. This guide follows the 105-slide deck and cites a page for every claim.

MIT 6.5940 Lecture 7: NAS I — From Hand-Designed Building Blocks to Search Spaces and Search Strategies

Lecture 7 has three parts. It first reviews fully connected, convolution, grouped, depthwise, and 1×1 convolution layers through their MAC formulas. It then takes apart how the ResNet bottleneck, ResNeXt, MobileNet, MobileNetV2, ShuffleNet, and the Transformer each save compute; the bottleneck, for example, needs 8.5× fewer MACs than a plain 3×3 convolution over 2048 channels. The last part is NAS: search spaces are either cell-level or network-level (depth, resolution, width, kernel size, topology), and there are five search strategies: grid, random, reinforcement learning, gradient descent, and evolution. One arithmetic exercise on the slides shows that the NASNet cell space already holds 3.2×10¹¹ candidates at M=5, N=2, B=5.