🌏 中文版
This is part 9 of the Reading MIT 6.5940 series. It covers Lecture 8: Neural Architecture Search (Part II) from the Fall 2024 course page, taught by Song Han on October 1, 2024. Both materials are public:
- Slides: Lec08-Neural-Architecture-Search-II.pdf (105 pages; page numbers below are PDF pages)
- Video: EfficientML.ai Lecture 8 - Neural Architecture Search Part II
Using the access grades from the course map, Fall 2024 is A3, enough for self-study. Fall 2026 comparison: as of 2026-09-30, the Fall 2026 course page has released only L1–L6, and the Lecture 8 slide and video links are still empty. This post uses Fall 2024 only.
Where the last lecture left off: how do you score a candidate?
The previous post covered two parts of NAS: the search space (the set of candidate architectures) and the search strategy (how to move through it). Slide 5 adds the third part, the accuracy estimation strategy: given an architecture, how do you estimate its accuracy? The spine of this lecture is that estimation keeps getting cheaper:
| Stage | Approach | Cost |
|---|---|---|
| Train from scratch | Fully train every candidate | Highest |
| Inherit weights / hypernetwork | Borrow weights from another model or a generator | Saves part of training |
| Once-for-All | Train one big network, extract subnets directly | One training run; seconds per evaluation |
| Zero-shot | No training, just compute a score | One forward/backward pass |
The lecture plan on slide 3 follows the same order: accuracy estimation → hardware-aware NAS → zero-shot NAS → neural-hardware architecture co-search → NAS applications.
1. Estimating accuracy
Training from scratch: too expensive beyond small datasets
Slide 7 cites Zoph and Le (ICLR 2017): 12,800 architectures trained on CIFAR-10, at a cost of 22,400 GPU-hours. Slide 8 then asks what happens on ImageNet or COCO. The answer is that you can't.
Inheriting weights: don't start from zero every time
Slide 10 uses Net2Net: a new architecture inherits weights from a parent model (Net2Wider widens it, Net2Deeper deepens it), which cuts training cost. Slide 11 goes one step further with Cai et al. (AAAI 2018). Instead of generating architectures, the controller generates network transformation actions, such as "make it wider" or "make it deeper," applied to an existing model.
Hypernetworks: one network generates weights for others
SMASH (slide 13) works in three steps. At each training step, sample a random architecture from the search space. Let the hypernetwork generate weights for that architecture. Update the hypernetwork by gradient descent. After training, any candidate can get a set of weights for evaluation.
2. Hardware-aware NAS: put the target hardware in the loop
Why search separately for each device
Slide 15 states the position plainly: one model for GPU, CPU, phone, and Raspberry Pi is inefficient; specialized models are efficient. The problem was cost, which pushed earlier NAS onto proxy tasks. Slide 16 gives two examples. NASNet needed 48,000 GPU hours, roughly five years on a single GPU, even on CIFAR. DARTS would need 100GB of GPU memory to search directly on ImageNet. So people searched on CIFAR-10, in smaller spaces, with fewer epochs, and used FLOPs and parameter counts as the efficiency metric. An architecture that wins on the proxy is not guaranteed to win on the real task and hardware.
ProxylessNAS: keep only one path alive
ProxylessNAS (slides 17–19):
- Build an over-parameterized network that holds every candidate path at each layer.
- Reduce NAS to a single training run of that network.
- Prune redundant paths based on architecture parameters.
The key is slide 18. Binarize the architecture parameters so that only one path's activations are in memory at a time. Memory drops from O(N) to O(1). That lets the search run directly on ImageNet, in a large space, with full training, and with measured latency instead of FLOPs as the efficiency signal (the comparison table on slide 19).
MACs are not real latency
Slides 20–22 carry the idea from this lecture most worth keeping. The slides use measurements from HAT:
- On an NVIDIA Titan Xp GPU, adding layers and widening the hidden dimension can reach similar FLOPs with very different latency (slide 21).
- Widening the hidden dimension has a large effect on Raspberry Pi ARM CPU latency and almost none on the GPU (slide 22).
The same MACs number means different latencies on different hardware. If you want the result to be fast on the target device, that device's latency has to be the feedback signal.
Getting latency: measure, look up, or predict
MnasNet measured every candidate on a phone. Slides 23–24 call that slow and expensive. ProxylessNAS (slides 25–30) instead builds an [architecture, latency] dataset and fits a latency model, in one of two forms:
- Layer-wise: a latency lookup table. Measure each op's latency on the device once, store it, and sum the entries for a candidate's layers (slides 26–29).
- Network-wise: a latency prediction model. Predict latency from features of the whole architecture. Slides 31–32 use HAT as the example, with features such as layer count, embedding dim, hidden dim, and head count. On a Raspberry Pi ARM CPU, predicted and measured latency fall almost exactly on the y=x line.
Slide 33 reports the payoff: the model specialized for mobile is 1.83x faster than the non-specialized one, and the gap is larger on GPU.
Once-for-All: train once, extract one per device
Even if each search trains only one big network, costs add up as devices multiply. Slides 37–39 illustrate this with a MnasNet-style search-train-retrain loop: design cost grows from 40K GPU hours to 160K, then to 1600K when repeated for many devices.
Once-for-All (OFA) (slides 40–73) replaces the loop:
- Train a once-for-all network whose subnets are sparsely activated parts of it.
- At deployment time, pick a subnet and get its accuracy and latency.
- Repeat step 2 and keep the best.
Slide 40 puts the two flows side by side. The old flow trains once per piece of feedback (about a day each); OFA evaluates an extracted subnet in seconds. Slide 46 says one OFA network holds about 10^19 subnets that share weights and are trained jointly, which amortizes the training cost. Slide 43 extends the target list down to three microcontrollers: STM32H743 (512kB SRAM / 2MB Flash), STM32F746 (320kB / 1MB), and STM32F412 (256kB / 1MB).
Progressive shrinking: training 10^19 subnets together without them fighting (slides 47–71)
OFA trains large-to-small, opening up four elastic dimensions one at a time:
| Dimension | How |
|---|---|
| Resolution | Sample a random input image size for each batch |
| Kernel size | Start with the full 7×7; a smaller kernel takes the centered weights and multiplies them by a transformation matrix (the slides show 25×25 and 9×9) |
| Depth | Train at full depth, then gradually allow later layers in each unit to be skipped |
| Width | Train at full width, then shrink gradually; before shrinking, sort channels by importance and keep the most important ones |
Slide 72 adds a hardware observation. On Xilinx ZU9EG and ZU3EG FPGAs, OFA-designed models have higher arithmetic intensity (Ops/Byte), so they are less memory-bound and reach higher utilization and GOPS/s without changing the RTL. This connects to the efficiency metrics in part 1 and to the peak-memory constraints in Lab 3.
3. Zero-shot NAS: skip training entirely
Slide 75 states the goal: estimate accuracy by analyzing the architecture, without training it. The slides give two methods.
Zen-NAS (slide 76):
- Draw a random input x ~ N(0,1) and perturb it to get x′ = x + ε.
- Initialize all network weights from N(0,1).
- Compute z₁ = log‖f(x′) − f(x)‖. The intuition: a good model should be sensitive to input perturbations.
- Add a batch-normalization variance term z₂ over the layers. The Zen score is z₁ + z₂.
GradSign (slide 77) starts from the intuition that a good model has denser sample-wise local minima, so gradients from different samples are more likely to share the same sign at initialization. That sign-agreement statistic becomes the score.
The slides do not state how well zero-shot scores correlate with trained accuracy, and this post doesn't fill that in.
4. Searching the network and the accelerator together: NAAS
Everything so far fixes the hardware and searches the model. NAAS, starting on slide 79, puts the accelerator into the search space too. Slide 80 lays out three layers:
| Layer | Searchable dimensions |
|---|---|
| Accelerator | local buffer size, global buffer size, #PEs, compute array size, PE connectivity |
| Compiler (mapping) | loop order, loop tiling size, dataflow |
| Neural network | #layers, #channels, kernel size, bypass, input/weight quantization precision |
Slide 81's claim: searching both in one optimization loop gives better-matched solutions.
There is a practical snag (slides 84–89). Parameters like loop order aren't numbers. Index-based encoding (CRXKYS as 0, CXYRSK as 1) is meaningless: adding or subtracting one from an index carries no physical information. NAAS uses importance-based encoding instead. Fix each dimension's position in the vector, let the optimizer assign each dimension a numerical importance, and sort by importance in decreasing order to get the loop order (or take the top two as the parallel dimensions).
Results (slides 91–92): compared with searching architectural sizing only, also searching connectivity and mapping gives considerably larger EDP (energy-delay product) reductions. NAAS plus OFA beats the baseline human design by +2.7% accuracy with 4.4x lower EDP. Slide 92 also notes that the dataflow NAAS found parallelizes output height and output channel, which is very different from the human design.
5. Applications: the OFA idea in other domains
Slides 94–101 are a tour, one slide per example:
- NLP: HAT. Slide 94: for WMT'14 En-Fr on a Raspberry Pi, compared with the Evolved Transformer, 2.7x faster, 3.7x smaller, 3.2x fewer FLOPs, 10,148x lower search cost, and 0.1 higher BLEU.
- Point clouds: SPVNAS. Slide 96: MinkowskiNet at 3.4 FPS versus SPVNAS at 9.1 FPS.
- GANs: Anycost GAN. Train once; use a small subnet for fast previews and a large one for the final high-quality result, aimed at interactive editing on an iPad (slide 97).
- Pose estimation: Lite Pose, on-device pose estimation via hardware-aware NAS (slide 98).
- Quantum circuits: QuantumNAS. Slide 100: quantum noise drops accuracy from 87% to 47%. The approach trains a "super circuit," searches for a noise-robust sub-circuit, and prunes small-magnitude gates. The topic returns in Lecture 23.
- LLMs: Flextron (ICML 2024). Same model, same weights; at inference a router picks how much of the MLP and attention to use for a latency target. Converting a trained LLM means ranking heads and channels, grouping them, and training the router (slide 101).
Where to go next
Slide 102 summarizes five items: performance estimation in NAS, hardware-aware NAS, zero-shot NAS, neural-hardware architecture search, and NAS applications.
These ideas land immediately in Lab 3. You get an OFA-trained MCUNetV2 super network, implement an efficiency predictor (MACs and peak memory) and an accuracy predictor, and write random and evolutionary search. The next lecture, L9 Knowledge Distillation, turns to a different question: once the architecture is fixed, how do you train a small model better?
If you have one hour: watch the ProxylessNAS-to-OFA stretch of the video (slides 16–73), which is what Lab 3 uses directly. Zero-shot NAS and NAAS can wait.
References
- MIT 6.5940 Fall 2024 course page
- MIT 6.5940 Fall 2026 course page
- Lec08-Neural-Architecture-Search-II.pdf (Fall 2024 slides)
- EfficientML.ai Lecture 8 - Neural Architecture Search Part II (YouTube)
- Zoph and Le, Neural Architecture Search with Reinforcement Learning (ICLR 2017)
- Chen et al., Net2Net (ICLR 2016)
- Cai et al., Efficient Architecture Search by Network Transformation (AAAI 2018)
- Brock et al., SMASH (ICLR 2018)
- Cai et al., ProxylessNAS (ICLR 2019)
- Tan et al., MnasNet (CVPR 2019)
- Wang et al., HAT (ACL 2020)
- Cai et al., Once-for-All (ICLR 2020)
- Lin et al., Zen-NAS (ICCV 2021)
- Zhang and Jia, GradSign (ICLR 2022)
- Lin et al., NAAS (DAC 2021)
- Tang et al., SPVNAS (ECCV 2020)
- Lin et al., Anycost GANs (CVPR 2021)
- Wang et al., Lite Pose (CVPR 2022)
- Wang et al., QuantumNAS (HPCA 2022)
- Cai et al., Flextron (ICML 2024)
- Global AI/CS course map (A0–A3 grades)
Loading...