Skip to content

MIT 6.5940 Lab 3: Finding a Microcontroller Model with a Supernet, Predictors, and Evolutionary Search

Sep 30, 20261 min
TL;DRLab 3 of MIT 6.5940 (Fall 2024) hands you an OFA-trained MCUNetV2 super network (more than 10^19 subnets) and the Visual Wake Words dataset. Ten questions, 100 points plus 10 bonus: implement a MACs/peak-memory efficiency predictor and a three-layer MLP accuracy predictor, write random search and evolutionary search, then find a subnet that reaches at least 92.5% accuracy under 250KB and 60M MACs. This guide maps the question structure and what each question trains. It does not include solutions.

🌏 中文版

This is part 10 of the Reading MIT 6.5940 series. It covers Lab 3: Neural Architecture Search from the Fall 2024 course page. The schedule lists it as released on October 8, 2024 (the day of Lecture 10, MCUNet) and due October 22 (the day of Lecture 13). The assignment is a public Colab notebook titled "MIT 6.5940 EfficientML.ai Lab 3: Neural Architecture Search".

It draws on the two lectures before it: the search space and search strategy from L7, and Once-for-All and efficiency prediction from L8. Read the L8 post before opening the notebook.

Fall 2026 comparison: the lab list on the Fall 2026 course page includes Lab 3 NAS, but it had not been released as of 2026-09-30. Whether it reuses this notebook is unknown.

What the lab asks you to do

The notebook states its goal in one sentence:

"In this lab, you will learn how to search for a tiny neural network that can run efficiently on a microcontroller."

The assignment has two main parts. Points come from the notebook:

SectionQuestionsPoints
Getting Started: super network and VWW datasetQ15
Part 1: PredictorsQ2–Q430
Part 2: Neural Architecture SearchQ5–Q1065 + 10 bonus

That is 100 points plus a 10-point bonus.

Before you start: environment and downloads

The first code cell installs graphviz, thop, and onnx, then downloads two archives from Dropbox: the MCUNet code (mcunetv2-dev-main.zip) and the VWW dataset (data.zip). I tested both links on 2026-10-01. They still worked, at about 31.7MB and 376MB.

Later cells move the network to cuda:0, so you need a GPU runtime. The notebook does not say which GPU is the minimum.

Getting Started: what the super network looks like (Q1, 5 points)

The notebook uses the MCUNetV2 super network, trained the Once-for-All way. The dataset is Visual Wake Words: binary classification sampled from COCO, asking whether a person is in the image.

The notebook specifies the design space:

DimensionChoices
Block typeinverted MobileNet block
Depthwise kernel size3, 5, 7
Expand ratio3, 4, 6
Depth per stagebase depth to base depth + 2
Global width multiplier0.5×, 0.75×, 1.0×
Input resolution (used by the predictor)96, 112, 128, 144, 160

The notebook says this OFAMCUNets contains more than 10^19 subnets. You can extract and evaluate a subnet without training it, and the expected accuracy is roughly 83.6–88.7%.

Watch the data split. In build_val_data_loader, split=0 is the real validation set and must not be used for architecture search. split=1 is a holdout minival set, used to build the accuracy dataset and to recalibrate BN statistics.

Q1 asks you to sample several subnets by hand, try different input resolutions, and write down what you see. The hint asks which dimension affects accuracy most.

Part 1: two predictors (Q2–Q4, 30 points)

Search evaluates many subnets, and running real inference on each one is too slow. The notebook cites OFA: measuring one subnet's accuracy on ImageNet takes about a minute, while a predictor takes under a second per subnet.

Q2 (10 points): efficiency predictor. Implement AnalyticalEfficiencyPredictor.get_efficiency. It uses hook-based analysis to return a subnet's MACs (in millions) and peak memory (in KB). The hints point to two existing functions, count_net_flops and count_peak_activation_size. The other method, satisfy_constraint, is already written: it compares each measured value against the constraint and returns False if any one is exceeded.

Q3 (10 points): accuracy predictor. A subnet is first turned into a binary vector by MCUNetArchEncoder: resolution, width multiplier, and each block's kernel size and expand ratio are one-hot encoded and concatenated. The notebook explains the choice of one-hot over numeric values: all of these hyperparameters are discrete. You implement a three-layer MLP with 400 channels in the hidden layers.

Q4 (10 points): train the accuracy predictor. The accuracy dataset has 50,000 [architecture, accuracy] pairs: 40,000 for training and 10,000 for validation. The training target is accuracy - base_acc, not raw accuracy (base_acc is the mean accuracy over the dataset). The difference is small and easier to learn. The loss is L1 and the optimizer is Adam, and the notebook says training takes about 1–2 minutes. For full credit, the scatter plot of predicted vs. true values must show a linear relationship.

Part 2: search (Q5–Q10, 65 points + 10 bonus)

Q5 (5 points): random search. RandomSearcher keeps sampling until it collects n_subnets subnets that meet the constraint, then scores them with the accuracy predictor. You fill in one line: pick the index with the highest score.

Q6 (5 points): search_and_measure_acc. You only add the line that calls the searcher. The rest of the code takes the best subnet, recalibrates BN on the holdout data, and measures accuracy on the real validation set. The notebook then runs four constraints: MACs ≤ 50M and ≤ 100M, peak memory ≤ 256KB and ≤ 512KB, with 300 samples each.

Something worth checking yourself: the random-search cell builds its constraint with the key millonMACs (missing an "i"), while the predictor returns millionMACs. satisfy_constraint skips any measured key that the constraint dict does not contain, so those two MACs constraints may not take effect at all. The evolutionary-search cell spells it correctly. If your MACs numbers look wrong, look here first.

Q7 (20 points): evolutionary search. The EvolutionSearcher skeleton is provided, with hyperparameters arch_mutate_prob, resolution_mutate_prob, population_size, max_time_budget, parent_ratio, and mutation_ratio. Mutation (resolution, width, and architecture each change with some probability) is already written. You fill in two places. The first is crossover: for each field and each block, pick the value from one of the two parents at random. The second is selection at the start of each generation: sort by predicted accuracy, highest first, and keep only the top parents_size as parents. The code that refills the population with mutation and crossover is already there.

Q8 (10 points): tune evo_params. The defaults are deliberately small (population 10, time budget 10). You raise them, compare results, and write up what you find.

Q9 (15 points + 10 bonus): real constraints. The notebook cites a TensorFlow blog post on VWW to explain that real applications often have several constraints at once, then sets two tasks:

  • 250KB and 60M MACs: accuracy ≥ 92.5% for full credit (15 points)
  • Bonus: 200KB and 30M MACs: accuracy ≥ 90% for full credit (10 points)

The hint says the two tasks do not need the same evo_params.

Q10 (10 points): feasibility. In the current design space, can you find a subnet that satisfies each of these?

  • A: activation at most 256KB and MACs at most 15M
  • B: activation at most 64KB

This one is not coding. You use the predictors from earlier to reason about the edges of the design space.

What you take away

  • You see every link in the L8 pipeline up close. The OFA super network means evaluating an architecture needs no training, and the predictor cuts evaluation further, from inference down to one MLP call.
  • The efficiency metrics become the two numbers an MCU cares about, MACs and peak memory (activations), rather than parameter count. That leads straight into the next two lectures: after L9 Knowledge Distillation, L10 MCUNet takes on KB-scale memory directly.
  • A feel for evolutionary search hyperparameters: how population size, mutation rates, and parent ratio change the accuracy you can reach under a constraint.

Limits for self-learners

  • No public solutions and no autograder. Submissions go through MIT's Canvas. Outside readers can rely only on the checks built into the notebook: Q2 compares the smallest and largest subnets against previously measured numbers, Q4 checks whether the scatter plot is linear, and Q9 measures accuracy on the validation set against thresholds stated in the question.
  • You need a GPU runtime. The code and dataset downloads live on Dropbox. The links worked on 2026-10-01 but may break later.
  • The short-answer questions (Q1, Q8, Q10) have no rubric. You can only check your reasoning against the L8 slides.

References