Diffusion is slow because one large network runs dozens to thousands of times, starting from pure noise. Lecture 18 first covers DDPM, conditioning, latent diffusion, SDEdit, and DreamBooth, then attacks the cost three ways: fewer steps (DDIM skips steps, progressive distillation halves the step count each round), less compute per step (DC-AE compresses images 64x, recomputing only the edited 1.7% region cuts MACs 8.2x, SVDQuant runs FLUX in 4-bit), and more devices (DistriFusion is up to 6.1x faster on 8 A100s).
Lecture 16 covers ViTs. At high resolution, attention cost grows with the square of the resolution. Window attention (Swin) confines computation to local windows, EfficientViT uses ReLU linear attention to get linear cost and then restores local and multi-scale ability, and SparseViT prunes unimportant windows. Self-supervised learning (contrastive learning, CLIP, MAE) answers the ViT's hunger for labeled data. HART pairs discrete tokens with residual diffusion and reaches several times the throughput of diffusion models. Lecture 17 targets three kinds of redundancy: 2D spatial in GANs (GAN Compression, AnyCost GAN, DiffAugment), temporal in video (TSM, temporal modeling at zero FLOPs), and 3D sparsity in point clouds (PVCNN, SPVCNN, BEVFusion). The Fall 2026 schedule drops Lecture 17.
Lecture 9 of MIT 6.5940 (Fall 2024) has five parts: what knowledge distillation (KD) is and why temperature matters; six things a student can match (logits, weights, features, gradients, sparsity patterns, relations); self and online distillation, which drop the fixed large teacher; KD for detection, segmentation, GANs, NLP, and LLMs; and Network Augmentation, built for tiny models. Raising the temperature from T=1 to T=10 moves the teacher's cat-vs-dog output from 0.982/0.017 to 0.599/0.401. That shift is where KD starts passing on dark knowledge.
Lecture 8 of MIT 6.5940 (Fall 2024) attacks the most expensive step in NAS: evaluating candidates. Training 12,800 architectures from scratch cost 22,400 GPU-hours, so the lecture walks through inherited weights, hypernetworks, ProxylessNAS's single-path training, latency lookup tables and predictors, Once-for-All's one training run for 10^19 subnets, training-free zero-shot NAS, and NAAS, which searches the network and the accelerator together. This guide follows the 105-slide deck and cites a page for every claim.