A week-by-week reading of Harvard CS181 (Machine Learning): the linear algebra, calculus, and probability refreshers, then linear regression and the models that follow, based on public homework.
CS181 2026 is A3 with hw0–6 as the weekly clock (no public recordings); 2025 adds a practical, 2024 has two midterms, 2023 was taught by Weiwei Pan. Start with HW0, then follow hw1→hw6.
HW0 checks CS181 prerequisites in four problems — y=Xw solvability, optimizing an objective, reasoning about randomness, and OLS in Python. The problem that slows you down most is the gap to patch before HW1.
HW1 uses an 800,000-year ice-core temperature dataset across four problems: kNN and kernel regression, a geometric proof of least squares, basis-function regression, and a probabilistic derivation of ridge and LASSO, ending with a coordinate-descent LASSO implementation.
CS181 Spring 2026 HW2 (due Feb 27) has four problems worth 90 points: train 10 logistic models on planet observations to see bias and variance, derive the MLE of a generative classifier, implement five classifiers on 27 loan applicants, and watch ridge reshape the loss surface under SGD, momentum, and Adam. The core skill is separating what one model's probability says from how much 10 models disagree.
CS181 Spring 2026 HW3 (due Mar 23) has three problems worth 100 points: unpack polynomial and RBF kernels into feature maps and go from ridge to dual coefficients α and the support-vector intuition; hand-derive backprop for a two-layer sigmoid network; and, for half the grade, train ResNets on Fashion-MNIST, measure your own scaling law, and use C≈6ND to split a fixed compute budget between model and data.
The CS181 Spring 2026 midterm is in class on Mar 10, worth 15% of the grade, closed-book with one double-sided note sheet. The official midterm checklist has four blocks: regression, classification, neural networks and model selection, and SVMs. This post maps each block to HW0–HW3 problem numbers, flags what the 2026 homework never drilled, and explains how to use the 2025 practice exam, review session, and concept checks.
HW4 Problem 1 (40 pts) takes self-attention apart in five steps: a 2×2 hand calculation, why we divide by √dk, permutation equivariance without positional encoding, single-head attention in pure NumPy, and multi-head attention in PyTorch with an attention heatmap on synthetic data.
HW4 Problem 2 has you train a convolutional autoencoder on 64×64 CelebA, sample from N(0, I), and watch it fail to produce faces. You then derive the ELBO, the reparameterization trick, and the closed-form KL, and turn the same backbone into a VAE to compare reconstructions and samples.
HW4 Problem 3 quantifies why voting trees get more accurate in three steps: with p=0.6 the Hoeffding bound needs B≈691 independent trees to reach 10⁻⁶; correlation ρ between trees floors ensemble variance at ρσ²; and random forests' dense ensembling is compared with MoE's sparse routing.
HW5 Problems 3–4 give handwritten digits to three methods that never see a label: K-means summarizes the data with 10 mean images, HAC builds a merge tree you can cut at any number of clusters, and PCA compresses images onto a few continuous directions. All three answer the same question — how much error do you pay to describe the data with a few objects — and the assignment makes you compare their objectives and reconstruction errors directly.
HW5 Problems 1–2 both turn learning into a classification task. SimCLR's NT-Xent loss asks the network to pick the other augmented view of the same image out of 2N−1 candidates; a GAN's discriminator classifies real versus fake. You derive the math behind each (the cross-entropy equivalence, the optimal discriminator and the JS divergence), then write both training loops on FashionMNIST and MNIST.
HW6 Problem 4 (20 points) takes apart the cost of generating one token at a time in three questions: picking the most likely token at each step doesn't give the most likely sequence; recomputing every key at every step makes cost quadratic, and a KV cache brings it back to linear; speculative decoding lets a small model guess and a large model verify in one pass. It is all pencil-and-paper, and every question maps onto a real design choice in today's LLM inference systems.
HW6 Problem 1 (15 pts) swaps the discrete HMM from lecture for a continuous state: the state drifts by Gaussian noise each step, each observation adds more noise, and you derive the mean and variance of the filtering distribution p(zₜ | x₀…xₜ). That is a one-dimensional Kalman filter. The solution is two moves, predict with the transition and then correct with the observation, and the problem hands you both Gaussian identities you need.
HW6 Problem 2 (15 pts) hands you a 4×5 Gridworld where moves can slip and rewards arrive only when you leave a cell. You write one step each of policy evaluation, policy iteration, and value iteration in the notebook, watch how the discount factor γ reshapes the policy, and finally ask whether this is a sensible model of the robot task at all. The rules of the world are fully known, so this is planning, not learning yet.
HW6 Problem 3 (20 pts) has you write a tabular Q-learning agent for Swingy Monkey, a Flappy Bird-like game. The minimum bar is scoring over 50 at least once within 100 epochs, plus one improvement of your choice. Problem 5 (10 pts) is a 250-word ethics question: assuming a social platform's users grow more politically extreme, use RL concepts to explain how the choice of reward function might have contributed.
The public materials for the CS181 2026 final (May 9) — the final checklist, 16 second-half practice problems, and a 66-page final review — are all 2025 versions. They cover Bayes nets and EM, which the 2026 schedule never lists, and skip Transformers, VAEs, GANs, and autoregressive models. Sort the checklist against the 2026 schedule into three buckets, cover the new topics with homework and sections, then do the 2025 practical for one end-to-end project.