🌏 中文版
This is part 10 of the Reading MIT 6.5940 series. It covers Lab 3: Neural Architecture Search from the Fall 2024 course page. The schedule lists it as released on October 8, 2024 (the day of Lecture 10, MCUNet) and due October 22 (the day of Lecture 13). The assignment is a public Colab notebook titled "MIT 6.5940 EfficientML.ai Lab 3: Neural Architecture Search".
It draws on the two lectures before it: the search space and search strategy from L7, and Once-for-All and efficiency prediction from L8. Read the L8 post before opening the notebook.
Fall 2026 comparison: the lab list on the Fall 2026 course page includes Lab 3 NAS, but it had not been released as of 2026-09-30. Whether it reuses this notebook is unknown.
What the lab asks you to do
The notebook states its goal in one sentence:
"In this lab, you will learn how to search for a tiny neural network that can run efficiently on a microcontroller."
The assignment has two main parts. Points come from the notebook:
| Section | Questions | Points |
|---|---|---|
| Getting Started: super network and VWW dataset | Q1 | 5 |
| Part 1: Predictors | Q2–Q4 | 30 |
| Part 2: Neural Architecture Search | Q5–Q10 | 65 + 10 bonus |
That is 100 points plus a 10-point bonus.
Before you start: environment and downloads
The first code cell installs graphviz, thop, and onnx, then downloads two archives from Dropbox: the MCUNet code (mcunetv2-dev-main.zip) and the VWW dataset (data.zip). I tested both links on 2026-10-01. They still worked, at about 31.7MB and 376MB.
Later cells move the network to cuda:0, so you need a GPU runtime. The notebook does not say which GPU is the minimum.
Getting Started: what the super network looks like (Q1, 5 points)
The notebook uses the MCUNetV2 super network, trained the Once-for-All way. The dataset is Visual Wake Words: binary classification sampled from COCO, asking whether a person is in the image.
The notebook specifies the design space:
| Dimension | Choices |
|---|---|
| Block type | inverted MobileNet block |
| Depthwise kernel size | 3, 5, 7 |
| Expand ratio | 3, 4, 6 |
| Depth per stage | base depth to base depth + 2 |
| Global width multiplier | 0.5×, 0.75×, 1.0× |
| Input resolution (used by the predictor) | 96, 112, 128, 144, 160 |
The notebook says this OFAMCUNets contains more than 10^19 subnets. You can extract and evaluate a subnet without training it, and the expected accuracy is roughly 83.6–88.7%.
Watch the data split. In build_val_data_loader, split=0 is the real validation set and must not be used for architecture search. split=1 is a holdout minival set, used to build the accuracy dataset and to recalibrate BN statistics.
Q1 asks you to sample several subnets by hand, try different input resolutions, and write down what you see. The hint asks which dimension affects accuracy most.
Part 1: two predictors (Q2–Q4, 30 points)
Search evaluates many subnets, and running real inference on each one is too slow. The notebook cites OFA: measuring one subnet's accuracy on ImageNet takes about a minute, while a predictor takes under a second per subnet.
Q2 (10 points): efficiency predictor. Implement AnalyticalEfficiencyPredictor.get_efficiency. It uses hook-based analysis to return a subnet's MACs (in millions) and peak memory (in KB). The hints point to two existing functions, count_net_flops and count_peak_activation_size. The other method, satisfy_constraint, is already written: it compares each measured value against the constraint and returns False if any one is exceeded.
Q3 (10 points): accuracy predictor. A subnet is first turned into a binary vector by MCUNetArchEncoder: resolution, width multiplier, and each block's kernel size and expand ratio are one-hot encoded and concatenated. The notebook explains the choice of one-hot over numeric values: all of these hyperparameters are discrete. You implement a three-layer MLP with 400 channels in the hidden layers.
Q4 (10 points): train the accuracy predictor. The accuracy dataset has 50,000 [architecture, accuracy] pairs: 40,000 for training and 10,000 for validation. The training target is accuracy - base_acc, not raw accuracy (base_acc is the mean accuracy over the dataset). The difference is small and easier to learn. The loss is L1 and the optimizer is Adam, and the notebook says training takes about 1–2 minutes. For full credit, the scatter plot of predicted vs. true values must show a linear relationship.
Part 2: search (Q5–Q10, 65 points + 10 bonus)
Q5 (5 points): random search. RandomSearcher keeps sampling until it collects n_subnets subnets that meet the constraint, then scores them with the accuracy predictor. You fill in one line: pick the index with the highest score.
Q6 (5 points): search_and_measure_acc. You only add the line that calls the searcher. The rest of the code takes the best subnet, recalibrates BN on the holdout data, and measures accuracy on the real validation set. The notebook then runs four constraints: MACs ≤ 50M and ≤ 100M, peak memory ≤ 256KB and ≤ 512KB, with 300 samples each.
Something worth checking yourself: the random-search cell builds its constraint with the key
millonMACs(missing an "i"), while the predictor returnsmillionMACs.satisfy_constraintskips any measured key that the constraint dict does not contain, so those two MACs constraints may not take effect at all. The evolutionary-search cell spells it correctly. If your MACs numbers look wrong, look here first.
Q7 (20 points): evolutionary search. The EvolutionSearcher skeleton is provided, with hyperparameters arch_mutate_prob, resolution_mutate_prob, population_size, max_time_budget, parent_ratio, and mutation_ratio. Mutation (resolution, width, and architecture each change with some probability) is already written. You fill in two places. The first is crossover: for each field and each block, pick the value from one of the two parents at random. The second is selection at the start of each generation: sort by predicted accuracy, highest first, and keep only the top parents_size as parents. The code that refills the population with mutation and crossover is already there.
Q8 (10 points): tune evo_params. The defaults are deliberately small (population 10, time budget 10). You raise them, compare results, and write up what you find.
Q9 (15 points + 10 bonus): real constraints. The notebook cites a TensorFlow blog post on VWW to explain that real applications often have several constraints at once, then sets two tasks:
- 250KB and 60M MACs: accuracy ≥ 92.5% for full credit (15 points)
- Bonus: 200KB and 30M MACs: accuracy ≥ 90% for full credit (10 points)
The hint says the two tasks do not need the same evo_params.
Q10 (10 points): feasibility. In the current design space, can you find a subnet that satisfies each of these?
- A: activation at most 256KB and MACs at most 15M
- B: activation at most 64KB
This one is not coding. You use the predictors from earlier to reason about the edges of the design space.
What you take away
- You see every link in the L8 pipeline up close. The OFA super network means evaluating an architecture needs no training, and the predictor cuts evaluation further, from inference down to one MLP call.
- The efficiency metrics become the two numbers an MCU cares about, MACs and peak memory (activations), rather than parameter count. That leads straight into the next two lectures: after L9 Knowledge Distillation, L10 MCUNet takes on KB-scale memory directly.
- A feel for evolutionary search hyperparameters: how population size, mutation rates, and parent ratio change the accuracy you can reach under a constraint.
Limits for self-learners
- No public solutions and no autograder. Submissions go through MIT's Canvas. Outside readers can rely only on the checks built into the notebook: Q2 compares the smallest and largest subnets against previously measured numbers, Q4 checks whether the scatter plot is linear, and Q9 measures accuracy on the validation set against thresholds stated in the question.
- You need a GPU runtime. The code and dataset downloads live on Dropbox. The links worked on 2026-10-01 but may break later.
- The short-answer questions (Q1, Q8, Q10) have no rubric. You can only check your reasoning against the L8 slides.
References
- MIT 6.5940 Fall 2024 course page
- MIT 6.5940 Fall 2026 course page
- Lab 3 Colab notebook (Fall 2024)
- Lin et al., MCUNetV2
- Lin et al., MCUNet (NeurIPS 2020)
- Cai et al., Once-for-All (ICLR 2020)
- Chowdhery et al., Visual Wake Words Dataset
- Visual Wake Words with TensorFlow Lite Micro (TensorFlow Blog)
- Global AI/CS course map (A0–A3 grades)
Loading...