Skip to content
All tags

#mit-65940

22 posts

MIT 6.5940 L1–L2 + Lab 0: How Do You Measure a Model's Size? Parameters, Activations, MACs, and Latency

The first two lectures of 6.5940 show that the problem exists, then hand you the rulers. L1 plots model parameter counts growing much faster than GPU memory, and contrasts 80GB on a cloud GPU with 320kB on a microcontroller. L2 splits efficiency metrics into memory metrics (#parameters, model size, peak activations) and compute metrics (MAC, FLOP, OP). AlexNet has 61M parameters and 724M MACs, and on a microcontroller the thing that runs out first is usually activation memory, not parameters. Lab 0 introduces a VGG variant on CIFAR-10 (9.2M parameters, 606M MACs) that later labs build on.

Reading MIT 6.5940: Song Han's Efficient AI Course Skipped a Year, So This Series Is Built on Fall 2024

MIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.

MIT 6.5940 L22–L23 Course Summary and Quantum Machine Learning: Pruning, NAS, and On-Device Training on Quantum Circuits

The last two lectures of MIT 6.5940 Fall 2024 come in two halves. The first half of Lecture 22 is a 13-page Course-Summary.pdf that redraws the course as three blocks (inference, training, application-specific) on System and Algorithm axes, then lays out the 7-item final project rubric. The second half, Quantum ML Part I, has a recording but no slides. Lecture 23 (Hanrui Wang, 99 slides) covers parameterized quantum circuits (PQCs): data encoding, parameter-shift gradients, probabilistic gradient pruning under noise (QOC), the TorchQuantum library, and QuantumNAS, which searches with a SuperCircuit and then prunes gates. It reads like a replay of the course's supernet and magnitude pruning on quantum circuits. Fall 2026 has replaced both lectures with a guest lecture.

MIT 6.5940 L18 Efficient Diffusion Models: Save on Steps, Resolution, and Compute per Step

Diffusion is slow because one large network runs dozens to thousands of times, starting from pure noise. Lecture 18 first covers DDPM, conditioning, latent diffusion, SDEdit, and DreamBooth, then attacks the cost three ways: fewer steps (DDIM skips steps, progressive distillation halves the step count each round), less compute per step (DC-AE compresses images 64x, recomputing only the edited 1.7% region cuts MACs 8.2x, SVDQuant runs FLUX in 4-bit), and more devices (DistriFusion is up to 6.1x faster on 8 A100s).

MIT 6.5940 L19–L20 Distributed Training: Split the Model When Memory Runs Out, Compress Gradients When Bandwidth Does

GPT-3's fp16 weights alone take 350GB, which does not fit on an 80GB A100, never mind gradients and Adam state. Lecture 19 covers how to split: data parallelism, ring all-reduce, ZeRO-1/2/3 (pushing the largest trainable model per 80GB GPU from 5B to 320B), GPipe raising pipeline utilization from 25% to 57%, Megatron-style tensor parallelism, and sequence parallelism with Ulysses and Ring Attention. Lecture 20 covers the communication bottleneck that follows: Alpa's automatic strategy search, DGC compressing gradients 277–608x without losing accuracy, TernGrad's three-value gradients, and DGA, which hides network latency behind delayed updates.

MIT 6.5940 L16–L17 Efficient Vision: What ViTs, GANs, Video, and Point Clouds Each Waste

Lecture 16 covers ViTs. At high resolution, attention cost grows with the square of the resolution. Window attention (Swin) confines computation to local windows, EfficientViT uses ReLU linear attention to get linear cost and then restores local and multi-scale ability, and SparseViT prunes unimportant windows. Self-supervised learning (contrastive learning, CLIP, MAE) answers the ViT's hunger for labeled data. HART pairs discrete tokens with residual diffusion and reaches several times the throughput of diffusion models. Lecture 17 targets three kinds of redundancy: 2D spatial in GANs (GAN Compression, AnyCost GAN, DiffAugment), temporal in video (TSM, temporal modeling at zero FLOPs), and 3D sparsity in point clouds (PVCNN, SPVCNN, BEVFusion). The Fall 2026 schedule drops Lecture 17.

MIT 6.5940 Fall 2026 Lab 1 Supplement: Reading GPU Bottlenecks with Roofline, the Profiler, and FlashAttention

This post covers Fall 2026 material, not the Fall 2024 edition the rest of the series follows. Fall 2026 replaced the pruning lab with "Efficient AI Fundamentals" (lab1_gpu_basics.zip). Part 1 has you hand-write a triple-loop GEMM and compute MAC, FLOPs, and I/O. Part 2 plots GEMM and GEMV rooflines. Part 3 works through a gemma-3-270m-it decoder layer, computing attention and MLP costs and comparing prefill with decode. Part 4 uses the PyTorch Profiler to inspect kernels, has you write GeLU to feel kernel fusion, then tries torch.compile and CUDA Graphs. Part 5 compares SDPA with FlashAttention. The core is 80 points plus 20 bonus, and all of Part 5 became bonus because Colab's T4 can't run it.

MIT 6.5940 L9 Knowledge Distillation: Teaching a Small Model Means Matching More Than Output Probabilities

Lecture 9 of MIT 6.5940 (Fall 2024) has five parts: what knowledge distillation (KD) is and why temperature matters; six things a student can match (logits, weights, features, gradients, sparsity patterns, relations); self and online distillation, which drop the fixed large teacher; KD for detection, segmentation, GANs, NLP, and LLMs; and Network Augmentation, built for tiny models. Raising the temperature from T=1 to T=10 moves the teacher's cat-vs-dog output from 0.982/0.017 to 0.599/0.401. That shift is where KD starts passing on dark knowledge.

MIT 6.5940 Lab 2: Implementing K-means and Linear Quantization, Down to an Integer-Only VGG

Lab 2 is a Colab notebook with 10 questions worth 100 points, built around a VGG pretrained on CIFAR-10. The first 3 questions cover K-means quantization: write the quantizer, work out how many clusters n bits gives you, write the centroid update, then compare accuracy at 8, 4, and 2 bits before and after fine-tuning. The other 7 cover linear quantization: write q = round(r/S) + Z, derive the scale and zero-point formulas, do per-channel weight quantization and bias quantization, write integer versions of the fully connected and convolution layers, and finally convert the whole model to INT8 for inference. This post lays out the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lab 3: Finding a Microcontroller Model with a Supernet, Predictors, and Evolutionary Search

Lab 3 of MIT 6.5940 (Fall 2024) hands you an OFA-trained MCUNetV2 super network (more than 10^19 subnets) and the Visual Wake Words dataset. Ten questions, 100 points plus 10 bonus: implement a MACs/peak-memory efficiency predictor and a three-layer MLP accuracy predictor, write random search and evolutionary search, then find a subnet that reaches at least 92.5% accuracy under 250KB and 60M MACs. This guide maps the question structure and what each question trains. It does not include solutions.

MIT 6.5940 Lab 4 + Lab 5: Quantizing an LLM with AWQ, Then Running LLaMA2-7B on Your Own Laptop

Lab 4 is a Colab notebook that rebuilds AWQ step by step on OPT-1.3B: first see how badly 3-bit quantization hurts perplexity, then keep 1% of the salient channels in FP16 (Q1), then protect them by scaling instead and search for the best scale (Q2). Each question is worth 50 points, plus a bonus scored on perplexity. Lab 5 moves to C++: run 4-bit LLaMA2-7B-chat on your own computer with TinyChatEngine and write five versions of the W4A8 linear-layer kernel (loop unrolling, multithreading, SIMD, multithreading plus unrolling, and all combined), 20 points each, plus up to 20 bonus points for performance. This post covers the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lecture 13: LLM Deployment Through Quantization, Sparsity, and Serving

Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.

MIT 6.5940 L14 LLM Post-Training: From SFT and RLHF to Fine-Tuning That Touches 1% of the Weights

Lecture 14 has three parts. Fine-tuning: SFT runs next-token prediction on desired answers, RLHF trains a reward model and then fine-tunes with KL-penalized RL, and DPO collapses both stages into one supervised step. Then comes a chain of PEFT methods: BitFit tunes only biases, Adapters add small layers but slow inference, Prompt/Prefix-Tuning eat input length, LoRA fixes latency with a low-rank branch you can merge back, QLoRA stores the backbone in NF4, and BitDelta compresses the fine-tune delta to 1 bit. Multimodal LLMs: Flamingo uses cross-attention, PaLM-E and VILA feed images in as tokens, and VILA-U can also output images. Prompt engineering: zero/few-shot, CoT, and RAG.

MIT 6.5940 L15 Long-Context LLM: When Context Grows, the KV Cache Breaks First

Lecture 15 has four parts. Extending context: interpolating RoPE stretches LLaMA from 2k to 32k, and LongLoRA's shifted sparse attention makes long-context fine-tuning cheap. Evaluation: lost-in-the-middle, Needle-in-a-Haystack, and LongBench. Efficient attention: the KV cache grows linearly with length. StreamingLLM finds that the first few tokens act as attention sinks, and keeping them plus a recent window gives stable generation. DuoAttention keeps a full KV cache only for a few retrieval heads. Quest keeps the whole KV cache but reads only the most critical pages for each query. The last part moves beyond Transformers: Mamba replaces attention with a selective SSM, and Jamba mixes the two.

MIT 6.5940 L10 MCUNet: Running Neural Networks on a Microcontroller with 320kB of SRAM

An MCU has roughly 256–320kB of SRAM and 1MB of Flash, tens of thousands of times less than a phone. Even an int8 MobileNetV2 needs 5x more peak memory than that. Lecture 10 answers with MCUNet: TinyNAS picks a search space before searching for a subnet, and MCUNetV2's patch-based inference cuts MobileNetV2's peak SRAM from 1372kB to 172kB. The lecture closes with tinyML applications in vision, audio, and anomaly detection.

MIT 6.5940 L8 NAS II: Scoring Architectures Without Training Them, and Putting Hardware in the Loop

Lecture 8 of MIT 6.5940 (Fall 2024) attacks the most expensive step in NAS: evaluating candidates. Training 12,800 architectures from scratch cost 22,400 GPU-hours, so the lecture walks through inherited weights, hypernetworks, ProxylessNAS's single-path training, latency lookup tables and predictors, Once-for-All's one training run for 10^19 subnets, training-free zero-shot NAS, and NAAS, which searches the network and the accelerator together. This guide follows the 105-slide deck and cites a page for every claim.

MIT 6.5940 Lecture 7: NAS I — From Hand-Designed Building Blocks to Search Spaces and Search Strategies

Lecture 7 has three parts. It first reviews fully connected, convolution, grouped, depthwise, and 1×1 convolution layers through their MAC formulas. It then takes apart how the ResNet bottleneck, ResNeXt, MobileNet, MobileNetV2, ShuffleNet, and the Transformer each save compute; the bottleneck, for example, needs 8.5× fewer MACs than a plain 3×3 convolution over 2048 channels. The last part is NAS: search spaces are either cell-level or network-level (depth, resolution, width, kernel size, topology), and there are five search strategies: grid, random, reinforcement learning, gradient descent, and evolution. One arithmetic exercise on the slides shows that the NASNet cell space already holds 3.2×10¹¹ candidates at M=5, N=2, B=5.

MIT 6.5940 L21 On-Device Training: Gradients Leak Data, and Activations Are the Memory Killer

There are two reasons to train on the device: the model has to adapt to each user's new data, and that data should not leave the device. Lecture 21 first shows that sharing only gradients is not safe either: Deep Leakage from Gradients recovers the original images and sentences from them. Then it tackles memory. Training costs more than inference because activations must be stored, not because of the parameters. TinyTL fine-tunes only biases plus a lightweight residual and saves 6.5x memory; SparseBP updates only the important layers and channels; QAS lets real int8 training match fp32; and PockEngine does autodiff at compile time, bringing training memory on a 256KB MCU down to 141KB.

MIT 6.5940 L3 Pruning I: Where to Prune, How Fine, and by What Criterion

Pruning removes unimportant weights or neurons from a neural network. The goal is written as minimizing loss subject to at most N nonzero weights. Lecture 3 of 6.5940 handles two of the decisions involved. First, granularity: from fine-grained pruning, which can remove any element, to channel pruning, which removes whole channels. The more regular the pattern, the easier it is to speed up on existing hardware, and the less you can remove. In between, 2:4 sparsity gives up to 2× speedup on NVIDIA Ampere GPUs. Second, criteria: look at weight magnitude, Batch Norm scaling factors, second derivatives, the fraction of zero activations, or how well a layer's output can be reconstructed after pruning.

MIT 6.5940 Lecture 6: Quantization II — PTQ Granularity and Clipping, QAT and STE, Binarization, Mixed Precision

Lecture 6 is about what to do when quantization costs you accuracy. First, without retraining: use finer scale granularity (per-channel, group, MX), clip outliers (EMA, calibration batches, MSE, KL), and round smarter (AdaRound). If that isn't enough, retrain: QAT keeps a full-precision copy of the weights, runs fake quantization in the forward pass, and uses the STE to pass gradients straight through. In the whitepaper table the slides cite, MobileNetV1 drops to 0.1% accuracy under per-tensor INT8 PTQ and recovers to 70.7% with per-channel QAT, against a 70.9% float baseline. The last two sections cover 1–2 bit binary and ternary networks, and HAQ, which uses reinforcement learning to assign a bit width to each layer.

MIT 6.5940 L11 TinyEngine and Parallel Computing: From Loop Tiling to In-Place Depthwise

Once algorithms have shrunk the model, how much more can the system layer squeeze out? Lecture 11 uses a single matrix multiply to show it: loop reordering gives 12x, tiling 19x (on an Intel Xeon 4114), and a CUDA version runs 94x faster end to end on a 2080Ti. The second half covers TinyEngine's inference tricks: im2col; in-place depthwise, which cuts peak memory from 2×C×H×W to (1+C)×H×W; NHWC for pointwise and NCHW for depthwise; and Winograd, with 2.25x fewer multiplications.

MIT 6.5940 L12 Transformer and LLM: The Architecture Seen Through an Efficiency Lens

In Lecture 12, 6.5940 switches from CNNs to Transformers. The lecture doesn't dwell on theory. It points to where memory and compute go. Attention is O(N²). If Llama-2-70B used MHA, its KV cache at batch 16 and length 4096 would take 160GB. GQA shrinks that 8x and MQA shrinks it 64x. MoE adds total parameters while keeping per-token compute flat. This post bridges into Lecture 13 on LLM deployment.