Homework 1 of CS224R Spring 2026 tests imitation learning on a custom Flappy Bird environment. The policy predicts 20 future target heights at once and executes only the first 10. You implement MSE-regression behavior cloning, a flow matching policy, and DAgger, then compare them in easy and hard modes. The PDF, LaTeX template, and starter code are all public, and a CPU is enough to run it. Solutions, the autograder, and Gradescope are not public. This guide covers what each problem asks you to build and answer. It does not give solutions.
Lecture 2 of CS224R Spring 2026 tackles two ways imitation learning fails. First, when demonstrations contain several reasonable behaviors, regression learns only their average. The fix is to make the policy a generative model (Gaussian mixtures, discretization plus autoregression, diffusion or flow matching) and to add action chunking. Second, compounding errors: once the policy slips, it reaches states the demonstrations never covered. The fix is DAgger or human-gated DAgger to collect corrections. The first two parts are exactly what HW1 covers.
MIT 6.S184 is a short IAP (January Independent Activities Period) course: five lectures (Lecture 3 is split into two recordings, 3-A and 3-B), three labs, and an 84-page set of lecture notes the course calls its backbone. Notes, slides, all six recordings, lab notebooks, and official solutions are public, so it grades A3, enough to self-study. Two gaps remain: lab submission goes through Gradescope inside Canvas, which only enrolled MIT students can use, and Lecture 5 on discrete diffusion has no lab.
Lab 2 turns §3–4 of the notes into PyTorch. You implement the Gaussian conditional path, its conditional vector field, and its conditional score. Then two nearly identical trainers do flow matching and score matching, Proposition 1 converts the learned vector field into a score, and finally a linear path makes a ring distribution flow into a checkerboard. Everything runs on 2D toy data. The README records a diffusion-coefficient bug fix dated 1/11/26; when I checked on 2026-09-30, the fix appeared only in the solutions notebook, not the student version, so patch it yourself before you start.
Lab 3 builds a conditional latent diffusion model on MNIST from scratch, in four stages: CFG training with label dropout (checked on a three-component Gaussian mixture), a diffusion transformer built piece by piece (Fourier time embedding, patchify, multi-head attention, adaLN-Zero, depatchify), a VAE, and finally the DiT trained inside the VAE's latent space. Problems and official solutions are public; submission goes through Gradescope on Canvas, which only enrolled MIT students can use.
Lecture 1 of MIT 6.S184 first rewrites "generate an image of a dog" as "sample from the data distribution," then gives the machine that does the sampling: start from Gaussian noise and simulate an ODE along a neural-network vector field (a flow model), or add a little Brownian-motion noise at every step to get an SDE (a diffusion model). Each is simulated with the simplest numerical method available, Euler and Euler–Maruyama. How to train the vector field is left to Lecture 2.
The object we want is the marginal vector field: run an ODE along it and noise flows into data. The catch is that it requires an integral over the whole dataset, so we can't compute it. Flow matching regresses on the conditional vector field instead, the one that pushes noise toward a single data point, which has a closed form. Theorem 12 in the notes shows the two losses differ by a constant and share the same gradient. On the CondOT path, training reduces to one line: sample data z, noise ε, and time t, and have the network predict z − ε at the point tz + (1−t)ε.
Feeding the prompt to the network as an extra input should, in theory, sample from p_data(x|y), but in practice the images don't follow the prompt closely enough. Lecture 3B uses Bayes' rule to split the guided vector field into the unguided vector field plus a classifier gradient; scaling that classifier term by w is classifier guidance. Replacing the classifier with the difference between guided and unguided fields gives CFG, which needs no classifier: ũ = (1−w)·u(x|∅) + w·u(x|y). Training only requires swapping the label for a null label ∅ with probability η. The costs: two network calls per step, and for w>1 you are no longer sampling from the data distribution.
The algorithms are complete by Lecture 3B; Lecture 4 tackles two engineering problems that show up at scale. First, the network must take an image, a time t, and a prompt and output a vector field of the same size, so we use a U-Net or a diffusion transformer (DiT), embedding time with Fourier features and text with frozen CLIP/T5 encoders. Second, pixel space is too big, so we first train a VAE to compress images into a latent space, run flow matching there, and decode at the end. Stable Diffusion 3 and Meta Movie Gen Video both follow this recipe: flow matching in latent space, a DiT variant, and CFG.
Text is a sequence of discrete tokens. There is no direction to move in, so ODEs and SDEs do not exist. Lecture 5 carries the recipe from Lectures 1–4 over unchanged and swaps only the underlying stochastic process: vector fields become rate matrices, ODEs become continuous-time Markov chains (CTMCs), and the continuity equation becomes the Kolmogorov forward equation. With the factorized mixture path, the marginal rate matrix has exactly one unknown: the probability of each position's original token given the noisy sequence. Training a discrete diffusion model therefore reduces to per-position classification with a cross-entropy loss. Make the noise all [mask] tokens and you get a masked diffusion language model.
HW9 has 19 questions worth 10 points, answered only on NTU COOL with no code submission. The first 16 cover four papers — DDPM, Flow Matching, Rectified Flow, and MeanFlow — ending with questions that compare their training signals and few-step generation. The last 3 require the Colab: train two small MLPs on a 2D Swiss roll, one Flow Matching model that learns instantaneous velocity (always evaluated with 50 Euler steps, converged at Histogram JS ≤ 0.10) and one MeanFlow model that learns average velocity (always one-step, ≤ 0.40). Then compare 1 step vs 1 step, Flow Matching across Euler step counts, and Euler vs RK4 at equal steps and at similar compute. The PDF includes a generative-modeling tutorial that skips most of the math, and every question is published in Chinese and English. Outside readers miss only the COOL grading and answers.