Skip to content
Series
10 posts

Reading MIT 6.S184

A lecture-by-lecture reading of MIT 6.S184 (IAP 2026) from the official lecture notes, slides, recordings, and three labs with solutions: ODEs/SDEs, flow matching, score matching, classifier-free guidance, DiT and latent spaces, and discrete diffusion.

Reading MIT 6.S184: Flow Matching and Diffusion Through ODEs and SDEs

MIT 6.S184 is a short IAP (January Independent Activities Period) course: five lectures (Lecture 3 is split into two recordings, 3-A and 3-B), three labs, and an 84-page set of lecture notes the course calls its backbone. Notes, slides, all six recordings, lab notebooks, and official solutions are public, so it grades A3, enough to self-study. Two gaps remain: lab submission goes through Gradescope inside Canvas, which only enrolled MIT students can use, and Lecture 5 on discrete diffusion has no lab.

MIT 6.S184 L1: Generation Is Sampling, and ODEs and SDEs Are the Machine

Lecture 1 of MIT 6.S184 first rewrites "generate an image of a dog" as "sample from the data distribution," then gives the machine that does the sampling: start from Gaussian noise and simulate an ODE along a neural-network vector field (a flow model), or add a little Brownian-motion noise at every step to get an SDE (a diffusion model). Each is simulated with the simplest numerical method available, Euler and Euler–Maruyama. How to train the vector field is left to Lecture 2.

MIT 6.S184 Lab 1: Simulating ODEs and SDEs

MIT 6.S184 Lab 1 has three parts. First you write the step functions for Euler and Euler–Maruyama. Then you use them to simulate Brownian motion and the Ornstein–Uhlenbeck process and watch how σ and θ shape trajectories and the final distribution. Finally you implement Langevin dynamics, watch a cloud of points get pushed toward a five-mode Gaussian mixture, and show by hand that the OU process is Langevin dynamics with a Gaussian target. The questions, code scaffolding, and official solutions are all on GitHub; outside readers get no Gradescope grading and have to check against the solutions themselves.

MIT 6.S184 L2: Flow Matching, Learning the Marginal Vector Field from Conditional Paths

The object we want is the marginal vector field: run an ODE along it and noise flows into data. The catch is that it requires an integral over the whole dataset, so we can't compute it. Flow matching regresses on the conditional vector field instead, the one that pushes noise toward a single data point, which has a closed form. Theorem 12 in the notes shows the two losses differ by a constant and share the same gradient. On the CondOT path, training reduces to one line: sample data z, noise ε, and time t, and have the network predict z − ε at the point tz + (1−t)ε.

MIT 6.S184 L3A: Score Functions, SDE Sampling, and Score Matching

A score function is the gradient of the log density; it points toward where probability rises fastest. On Gaussian paths, the score and last lecture's vector field are both linear in x and z, so each converts into the other (Proposition 1 in the notes): learn one and you have learned both. With the score in hand, you can add noise of any strength to the ODE and turn it into an SDE without changing the distribution at any time (Theorem 17). The score itself is learned with the same trick as flow matching, by regressing on the conditional score. On Gaussian paths, that amounts to predicting the noise that was added, which is the DDPM training objective.

MIT 6.S184 Lab 2: Writing Flow Matching and Score Matching by Hand

Lab 2 turns §3–4 of the notes into PyTorch. You implement the Gaussian conditional path, its conditional vector field, and its conditional score. Then two nearly identical trainers do flow matching and score matching, Proposition 1 converts the learned vector field into a score, and finally a linear path makes a ring distribution flow into a checkerboard. Everything runs on 2D toy data. The README records a diffusion-coefficient bug fix dated 1/11/26; when I checked on 2026-09-30, the fix appeared only in the solutions notebook, not the student version, so patch it yourself before you start.

MIT 6.S184 L3B: Guidance and Classifier-Free Guidance

Feeding the prompt to the network as an extra input should, in theory, sample from p_data(x|y), but in practice the images don't follow the prompt closely enough. Lecture 3B uses Bayes' rule to split the guided vector field into the unguided vector field plus a classifier gradient; scaling that classifier term by w is classifier guidance. Replacing the classifier with the difference between guided and unguided fields gives CFG, which needs no classifier: ũ = (1−w)·u(x|∅) + w·u(x|y). Training only requires swapping the label for a null label ∅ with probability η. The costs: two network calls per step, and for w>1 you are no longer sampling from the data distribution.

MIT 6.S184 L4: U-Nets, DiTs, and Latent Space

The algorithms are complete by Lecture 3B; Lecture 4 tackles two engineering problems that show up at scale. First, the network must take an image, a time t, and a prompt and output a vector field of the same size, so we use a U-Net or a diffusion transformer (DiT), embedding time with Fourier features and text with frozen CLIP/T5 encoders. Second, pixel space is too big, so we first train a VAE to compress images into a latent space, run flow matching there, and decode at the end. Stable Diffusion 3 and Meta Movie Gen Video both follow this recipe: flow matching in latent space, a DiT variant, and CFG.

MIT 6.S184 Lab 3: From DiT and VAE to Latent Diffusion

Lab 3 builds a conditional latent diffusion model on MNIST from scratch, in four stages: CFG training with label dropout (checked on a three-component Gaussian mixture), a diffusion transformer built piece by piece (Fourier time embedding, patchify, multi-head attention, adaLN-Zero, depatchify), a VAE, and finally the DiT trained inside the VAE's latent space. Problems and official solutions are public; submission goes through Gradescope on Canvas, which only enrolled MIT students can use.

MIT 6.S184 L5: Discrete Diffusion, Generating Language with CTMCs

Text is a sequence of discrete tokens. There is no direction to move in, so ODEs and SDEs do not exist. Lecture 5 carries the recipe from Lectures 1–4 over unchanged and swaps only the underlying stochastic process: vector fields become rate matrices, ODEs become continuous-time Markov chains (CTMCs), and the continuity equation becomes the Kolmogorov forward equation. With the factorized mixture path, the marginal rate matrix has exactly one unknown: the probability of each position's original token given the noisy sequence. Training a discrete diffusion model therefore reduces to per-position classification with a cross-entropy loss. Make the noise all [mask] tokens and you get a masked diffusion language model.