🌏 中文版
This is post 1 of the Reading MIT 6.5940 series, based on the Fall 2024 edition. The series entry point explains why it does not follow Fall 2026.
Official materials covered here:
- Lecture 1 Introduction: slides (93 pages), video
- Lecture 2 Basics of Neural Networks: slides (77 pages), video
- Lab 0: PyTorch Tutorial (Colab)
Page numbers below are PDF page numbers. The number printed in the slide corner is sometimes off by one or two.
The problem: models grow faster than hardware
Lecture 1 opens (page 3) with a two-line chart. One line is language-model size: Transformer 0.05B, BERT 0.34B, GPT-2 1.5B, GPT-3 175B, MT-NLG 530B. The other is GPU memory, from 32GB on the V100 to 80GB on the A100. The gap keeps widening, and the chart's caption reads "Model compression bridges the gap." Page 3 of Lecture 2 puts it in one line: Moore's law gives about 2× every two years, while deep learning models grow about 4× every two years.
Outside the cloud the gap is larger. Page 78 of Lecture 1 compares three platforms:
| Platform | Activation memory | Weight storage |
|---|---|---|
| Cloud AI | 80GB | ~TB/PB |
| Mobile AI | 4GB | 256GB |
| Tiny AI (microcontroller) | 320kB | 1MB |
Cloud GPUs and microcontrollers differ by five orders of magnitude in memory. Everything the course covers later (pruning, quantization, NAS, MCUNet) attacks the same question: how do you fit a model into the bottom two rows?
The middle of Lecture 1 walks through many HAN Lab projects: image recognition on phones, person detection on microcontrollers (MCUNet), EfficientViT-SAM, GAN Compression, TinyChat, and LLaMA-2 running on Jetson Orin with AWQ quantization. Each gets its own lecture later, so treat this part as a preview. The last page (page 93) lists the course goals: learn the key efficiency metrics of deep learning computation, accelerate inference and training on resource-constrained platforms, understand the trade-offs between optimization techniques, and deploy an LLM on your own laptop.
Four items in Lecture 2
The Lecture Plan on page 5 of Lecture 2 has four items:
- Review neural network terms: neuron, synapse, activation, feature, weight, parameter
- Review common layers: fully-connected, convolution, grouped convolution, depthwise convolution, pooling, normalization, transformer
- Introduce efficiency metrics: #Parameters, Model Size, Peak #Activations, MAC, FLOP, FLOPS, OP, OPS, Latency, Throughput
- Lab 0: PyTorch tutorial
One point from the terminology section matters later. When the course says "prune a synapse" it means pruning a weight. "Prune a neuron" means removing a whole output channel. Both mappings come up constantly in the pruning lectures.
Transformers get only two pages here (pages 39–40). The slides say the full architecture is covered in Lecture 12.
Layer shapes drive everything
Every metric below is computed from tensor shapes, so learn the notation first. The slides use n for batch size, cᵢ and cₒ for input and output channels, hᵢ/wᵢ and hₒ/wₒ for input and output height and width, k_h and k_w for kernel height and width, and g for the number of groups.
| Layer | Weight shape | Notes |
|---|---|---|
| Fully-connected | (cₒ, cᵢ) | Every output connects to every input |
| 2D Convolution | (cₒ, cᵢ, k_h, k_w) | Each output connects only to inputs in its receptive field; weights are shared |
| Grouped Convolution | (g·cₒ/g, cᵢ/g, k_h, k_w) | Channels split into g groups, each convolved separately with a narrower kernel |
| Depthwise Convolution | (c, k_h, k_w) | g = cᵢ = cₒ; one independent filter per channel |
| Pooling | none | No learnable parameters; stride usually equals kernel size |
The conv output size is hₒ = (hᵢ + 2p − k_h) / s + 1, where p is padding and s is stride (page 33). Page 32 gives the receptive field: L conv layers with kernel size k see L·(k − 1) + 1 pixels. A large image therefore needs many layers before any unit "sees" the whole picture, which is why networks downsample internally.
The normalization page (page 37) draws Batch Norm, Layer Norm, Instance Norm, and Group Norm with the same formula. They differ only in which set of pixels the mean and standard deviation are computed over. The Batch Norm scaling factor γ comes back in the pruning lectures.
Efficiency metrics: two families
Page 42 splits efficiency metrics into two groups:
- Memory-related: #parameters, model size, total/peak #activations
- Computation-related: MAC, FLOP/FLOPS, OP/OPS
Together they determine latency and energy. Let's take them in order.
#Parameters and model size
The parameter count is the number of elements in the weight tensors (ignoring bias):
| Layer | #Parameters |
|---|---|
| Linear | cₒ·cᵢ |
| Convolution | cₒ·cᵢ·k_h·k_w |
| Grouped Convolution | cₒ·cᵢ·k_h·k_w / g |
| Depthwise Convolution | cₒ·k_h·k_w |
The slides work through AlexNet layer by layer (page 56) and arrive at about 61M parameters. The largest layer is the first fully-connected layer: 4096 × (256×6×6) = 37,748,736. That single layer holds about 60% of the total.
Model size is the storage needed for the weights. When every weight uses the same data type, model size = #Parameters × bit width (page 58). AlexNet in 32-bit takes about 244MB; in 8-bit it drops to 61MB. That is the starting point for quantization in Lectures 5–6.
#Activations: often the real bottleneck
The title of page 60 says it directly: "#Activation is the memory bottleneck in CNN inference, not #Parameters."
The slides cite a comparison from MCUNet. ResNet-18 and MobileNetV2-0.75 both reach about 70% ImageNet top-1 accuracy. MobileNetV2 has 4.6× fewer parameters, but its peak activation does not shrink with them (page 61 is titled "#Activation didn't improve from ResNet to MobileNet-v2"). Page 62 plots memory use per MobileNetV2 block. The peak is 1372kB against a 256kB microcontroller limit, and it sits in the first few blocks.
Training makes this worse. Page 63 cites TinyTL: moving from ResNet-50 to MobileNetV2-1.4 cuts parameters by 4.3× but activation memory by only 1.1×.
The AlexNet example (page 65) shows two ways to count:
- Total #activations: the sum of every layer's output features. For AlexNet, 932,264
- Peak #activations: roughly one layer's input plus output. AlexNet peaks at the first conv layer: input 3×224×224 = 150,528 plus output 96×55×55 = 290,400, for 440,928
AlexNet's parameters live mostly in the fully-connected layers, while its activations pile up in the early conv layers. The two metrics blow up in different places. If you remember one thing from this post, make it that.
MAC, FLOP, OP
One MAC is a ← a + b·c. A matrix-vector product costs m·n MACs; a matrix-matrix product costs m·n·k (page 67).
MAC formulas per layer (batch size 1, bias ignored)
| Layer | MACs |
|---|---|
| Linear | cₒ·cᵢ |
| Convolution | cᵢ·k_h·k_w·hₒ·wₒ·cₒ |
| Grouped Convolution | cᵢ/g·k_h·k_w·hₒ·wₒ·cₒ |
| Depthwise Convolution | k_h·k_w·hₒ·wₒ·cₒ |
A conv layer's MAC count is its parameter count times the output area hₒ·wₒ, because the same weights are reused at every output position.
AlexNet totals 724M MACs (page 72). Compare that with the parameter counts. The first conv layer has only 96×3×11×11 = 34,848 parameters (page 56 prints 24,848, but the multiplication gives 34,848), yet it performs 105,415,200 MACs. The first fully-connected layer has 37.7M parameters and also 37.7M MACs. Conv layers have few parameters and lots of compute; fully-connected layers are the reverse.
The next three terms are the easiest to mix up:
- FLOP: a multiplication counts as one floating-point operation and an addition counts as another, so 1 MAC = 2 FLOPs. AlexNet has about 724M × 2 = 1.4G FLOPs (page 74)
- FLOPS: floating-point operations per second, FLOPS = FLOPs / second. This is hardware speed, not model compute
- OP/OPS: activations and weights are not always floating point (after quantization they may be integers), so OP is the generic operation count and OPS is operations per second (page 75)
FLOP and FLOPS differ by one letter: one describes a model, the other describes hardware. Papers and spec sheets often use them loosely, so read the context.
Latency and throughput
Latency is the delay to finish one task; throughput is how much data you process per unit time. Page 45 contrasts two designs:
| Latency | Throughput | |
|---|---|---|
| Design 1 | 50 ms | 20 image/s |
| Design 2 | 100 ms | 40 image/s |
Design 2 has higher latency and also higher throughput. The slides pose two questions: does high throughput imply low latency, and does low latency imply high throughput? Neither is guaranteed. Parallel processing can raise throughput without making any single task faster.
Page 46 gives an estimate:
Latency ≈ max(T_computation, T_memory)
T_computation is roughly the model's operation count divided by the processor's operations per second. T_memory is the time to move weights and activations: model size and activation size divided by memory bandwidth. Taking the max assumes compute and data movement overlap; the Fall 2026 slides draw this as a timeline on page 51. The formula ties the earlier metrics together. Parameters and activations set T_memory, and MACs set T_computation.
Energy: moving data costs more than computing
Page 47 cites Horowitz's ISSCC 2014 numbers for a 45nm process: a 32-bit integer add costs 0.1 pJ, a 32-bit float multiply 3.7 pJ, a 32-bit SRAM cache read 5 pJ, and a 32-bit DRAM read 640 pJ. One DRAM access costs thousands of times more than an integer add. Compression is not only about fitting in memory: every DRAM access you avoid saves energy.
Lab 0: the starting point for later labs
Lab 0 is a Colab notebook in six sections: Setup, Data, Model, Optimization, Training, Visualization.
- Data: CIFAR-10, 10 classes of 3×32×32 color images, batch size 512
- Model: a VGG-11 variant (fewer downsamples, smaller classifier). The backbone is 8 conv-bn-relu blocks with 4 max-pool layers in between
- Efficiency check: model size estimated from the parameter count, MACs counted with TorchProfile. The notebook states that the model has 9.2M parameters and needs 606M MACs per inference, and says the next few labs will work together on making it more efficient
- Training: cross-entropy loss, SGD with momentum, a custom learning-rate scheduler. It takes about 10 minutes and, if all goes well, reaches over 92.5% accuracy
Lab 0 itself is easy. What matters is getting to know this model. The Fall 2024 Lab 1 prunes the same VGG on CIFAR-10.
Fall 2026 comparison
The Fall 2026 L1 slides (91 pages, video) have the same structure as Fall 2024. The differences are in the logistics pages at the end; lab and grading changes are covered in the series entry point.
The L2 slides (86 pages, video) add three things:
- One more Lecture Plan item: "Review convolutional neural networks' architecture: AlexNet, VGG-16, ResNet-50, MobileNetV2," on pages 40–44. The ResNet-50 page draws the bottleneck block (1×1 → 3×3 → 1×1, with N/4 channels in the middle). The MobileNetV2 page draws the inverted bottleneck (1×1 expanding to N×6, then 3×3 depthwise, then 1×1 back to N)
- A timeline on the latency page (page 51), showing load input, load weight, compute, and store output overlapping, which is why the estimate takes the max
- Three new closing pages titled "Today's AI is too BIG" (pages 81–83). The model-vs-GPU-memory chart on page 81 extends to 2026
The Fall 2026 Lab 0 is nearly identical to Fall 2024 but adds two questions: Question 1.1 asks you to complete the model's forward pass, and Question 1.2 asks for the model's best accuracy. It also starts by mounting Google Drive and switching into the lab folder.
What you can do tonight
- Open the Lab 0 Colab, run through the Model section, and check that the printed parameter and MAC counts match the notebook's 9.2M and 606M.
- Work out the parameters and MACs of AlexNet's first conv layer by hand (96 filters of 11×11, 3 input channels, 55×55 output), then check against pages 56 and 72. Once you can do this, you will know how to turn a pruning ratio into compute saved in the next lecture.
Further reading
- MIT 6.7960 guide and CMU 11-785 guide: the full story of backpropagation and CNNs
- Stanford CS231N guide: how CNN architectures evolved
- Stanford CS336: GPUs and TPUs: latency from the memory-bandwidth angle
Series navigation: previous Series entry point | next Pruning I: granularity and criteria
References
- MIT 6.5940 Fall 2024 course page
- Lecture 1 slides: Introduction (Fall 2024)
- Lecture 1 video (Fall 2024)
- Lecture 2 slides: Basics of Neural Networks (Fall 2024)
- Lecture 2 video (Fall 2024)
- Lab 0: PyTorch Tutorial (Fall 2024, Colab)
- MIT 6.5940 Fall 2026 course page
- Lecture 2 slides (Fall 2026)
- Lab 0 (Fall 2026, Colab)
- Lin et al. (2020). MCUNet: Tiny Deep Learning on IoT Devices. NeurIPS
- Cai et al. (2020). TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learning. NeurIPS
- TorchProfile
Loading...