Skip to content

CS231N L14: Generative Models II — Why Adding Noise and Removing It Generates Images

Sep 30, 20261 min
TL;DRThe CS231N Spring 2026 diffusion lecture doesn't start with DDPM math. It opens by warning that terminology and notation in this area are a mess, then teaches one clean modern version: rectified flow. In training, pick a point between a data sample and noise and have the network predict the velocity from data toward noise; to generate, start from noise and walk backward for about 50 steps. The lecture then stacks on the practical pieces: classifier-free guidance, noise schedules that emphasize middle noise levels, diffusion on VAE latents, Transformers (DiT) as the denoiser, and distillation to cut the step count. Only at the end does it fold VP, VE, and ε/v-prediction into a generalized diffusion framework and name three mathematical views: latent variable model, score function, and SDE.

🌏 中文版

Source years: slides and assignments are from Spring 2026; the recordings are from Spring 2025 (YouTube). The two may differ. This post follows the 2026 slides and uses the recording only as a supplement.

This is part 16 of the Reading Stanford CS231N series. The previous post is L13: Generative Models I, Autoregressive Models, VAEs, and GANs; the next is L16: Vision and Language.

For the CS231N lecture on May 19, 2026, the schedule lists a single topic: Diffusion models. The official material is the 122-page lecture_14.pdf; the matching public recording is Spring 2025 Lecture 14.

The first 35 slides are actually about GANs, which the previous post already covers. This post starts at slide 36, where diffusion begins.

Scene: a field where "the notation is a mess"

Slide 36 lists five foundational diffusion papers: Sohl-Dickstein et al. 2015, Song & Ermon 2019, Ho et al. 2020 (DDPM), Song et al. 2021 (SDE), and Song et al. 2021 (DDIM).

Slide 37 follows with a warning, the line most worth remembering from this lecture: terminology and notation in this area are a mess. There are many mathematical formalisms and a lot of variance in terms and notation between papers. So the course makes a choice: teach only one modern, "clean" implementation — rectified flow.

That choice shapes how other material will feel. If you've read the DDPM paper first, the notation, the direction of time, and the prediction target here will all look different; the "generalized diffusion" section at the end of the lecture is what connects them.

Intuition: learn to remove a little noise, then repeat many times

Slide 41 puts diffusion in three sentences:

  1. Pick a noise distribution, usually a unit Gaussian
  2. Corrupt data x with varying noise levels t to get x_t; t = 0 is no noise, t = 1 is full noise
  3. Train a neural network f_θ(x_t, t) whose only job is to remove a little bit of noise

To generate, start from pure noise x₁ and apply f_θ many times in sequence to reach a noiseless sample x₀.

In the previous lecture's taxonomy, this is the "implicit density, iterative procedure to approximate samples" box: the model neither writes down p(x) nor generates in one shot like a GAN; it takes many small steps.

Mechanism 1: rectified flow training and sampling

Training (slide 45, based on Liu et al. 2022 and Lipman et al. 2022, Flow Matching): on each iteration, sample three things — noise z, data x, and t from a uniform distribution. Take the point x_t on the line between them; the target velocity v is the vector from data to noise, z − x. The network's job is to predict v, with an L2 loss.

Sampling (slide 56): choose a number of steps T (often T = 50), sample x from noise, and walk t from 1 to 0, computing v_t at each step and stepping 1/T in the opposite direction.

Slide 46 says the core training loop is "just a few lines of code". The code on the slide is an image, so the sketch below is my rewrite of the steps on slides 45 and 56, not the slide's code:

# Training: one iteration (rewritten from the steps on slide 45)
x = sample_data()                 # x ~ p_data
z = torch.randn_like(x)           # z ~ p_noise
t = torch.rand(x.shape[0])        # t ~ Uniform(0, 1)
x_t = (1 - t) * x + t * z
v = z - x
loss = ((f_theta(x_t, t) - v) ** 2).mean()

# Sampling (rewritten from the steps on slide 56)
x = torch.randn(shape)            # start from pure noise
for t in torch.linspace(1, 0, T + 1)[:-1]:
    v_t = f_theta(x, t)
    x = x - v_t / T
Formulas: rectified flow (slides 45, 56)

$$z \sim p_{noise},\quad x \sim p_{data},\quad t \sim \mathrm{Uniform}(0, 1)$$

$$x_t = (1 - t),x + t,z,\qquad v = z - x$$

$$\mathcal{L} = \big| f_\theta(x_t, t) - v \big|_2^2$$

At sampling time, t runs through 1, 1 − 1/T, 1 − 2/T, …, 0, and each step sets x ← x − f_θ(x_t, t)/T.

Mechanism 2: making generation follow instructions — conditioning and CFG

Conditional rectified flow (slide 61): feed a condition y (a class or text) to the network during training, and at generation time you can target p_data(x | y). The slides then ask: can we control how much we "emphasize" the condition y?

Classifier-free guidance (slides 67–71, Ho & Salimans): randomly drop y during training so the same model is both conditional and unconditional. For a noisy x_t:

  • The unconditional velocity v_∅ points toward p(x)
  • The conditional velocity v_y points toward p(x | y)
  • Their combination v_cfg points more strongly toward p(x | y)

Sampling steps along v_cfg. The "classifier-free" in the name contrasts with earlier work: Dhariwal & Nichol used a separate discriminative model p(y | x) and took the gradient of log p(y | x) with respect to x as the step direction. The slides' verdict on CFG: used everywhere in practice and very important for high-quality outputs, at the cost of doubling sampling time.

Formula: CFG (slide 67)

$$v^{\varnothing} = f_\theta(x_t, y_\varnothing, t),\qquad v^{y} = f_\theta(x_t, y, t)$$

$$v^{cfg} = (1 + w),v^{y} - w,v^{\varnothing}$$

Larger w puts more emphasis on the condition y.

Mechanism 3: which noise level is hardest?

Slide 78 asks: what is the optimal prediction for the network? Many (x, z) pairs can produce the same x_t, so the network must average over them.

  • Full noise (t = 1) is easy: the optimal v is the mean of p_data
  • No noise (t = 0) is easy: the optimal v is the mean of p_noise
  • Middle noise is the hardest and most ambiguous

Sampling t uniformly gives every noise level equal weight. The fix is a non-uniform noise schedule (slide 81): put more emphasis on middle noise, commonly with logit-normal sampling; for high-resolution data, shift toward higher noise to account for correlations between pixels. The citation here is Esser et al. 2024.

Slide 83 sums up: rectified flow is a simple, scalable setup for many generative modeling problems, but it doesn't work naively on high-resolution data. That leads to latent diffusion.

Mechanism 4: the modern pipeline — VAE + GAN + diffusion

Latent diffusion (slides 85–96, Rombach et al., CVPR 2022) has two stages:

  1. Train an encoder and decoder that compress an H×W×3 image into an H/D×W/D×C latent. A common setting is D = 8, C = 16, so a 256×256×3 image becomes 32×32×16; the encoder and decoder are CNNs with attention
  2. Freeze the encoder and train a diffusion model to denoise latents. To generate, start from a random latent, denoise it iteratively, and run the decoder to get an image

The slides call latent diffusion the most common form today.

How is the encoder/decoder trained? Both models from the previous lecture come back. It's a VAE, typically with a very small KL prior weight; the problem is that decoder outputs are often blurry. So add a discriminator, the GAN move. Slide 96 concludes: modern LDM pipelines use VAE + GAN + diffusion.

Diffusion Transformer (slide 99, Peebles & Xie, ICCV 2023): diffusion uses standard Transformer blocks, and the main question is how to inject conditioning. The diffusion timestep t most commonly goes in by predicting scale and shift; text, image, and similar conditions usually go in through cross-attention or joint attention.

A text-to-image example (slide 101) uses FLUX.1 [dev]:

ComponentSetting on the slide
Text encoderT5 + CLIP
Encoder/decoder8×8 downsampling
Diffusion model12B parameters
Image tokens64×64 = 1024 after 2×2 patchify

The output is a 1024×1024×3 image, and the latents being denoised are 128×128×16. Slides 104–105 extend the same architecture to text-to-video and list a long line of video diffusion models from 2024–2025 (Sora, Veo 2, MovieGen, Wan, Hunyuan, and more).

Distillation (slide 107): rectified flow sampling runs the model about 30–50 times, which is slow. Distillation algorithms reduce the number of steps, sometimes all the way to one, and can also bake CFG into the model.

Mechanism 5: one framework for everyone's notation

Only now does the lecture return to the "notation is a mess" warning from slide 37. Slides 110–114 replace each fixed coefficient in rectified flow with a function:

Formulas: generalized diffusion (slides 110–113)

$$x_t = a(t),x + b(t),z,\qquad y_{gt} = c(t),x + d(t),z,\qquad \mathcal{L} = | y_{gt} - f_\theta(x_t, t) |_2^2$$

  • Rectified flow: a(t) = 1 − t, b(t) = t, c(t) = −1, d(t) = 1
  • Variance Preserving (VP): a(t) = √σ(t), b(t) = √(1 − σ(t)); if x and z are independent with variance 1, x_t also has variance 1
  • Variance Exploding (VE): a(t) = 1, b(t) = σ(t); σ(1) must be large enough to drown out all signal in x
  • Prediction targets: x-prediction (c = 1, d = 0), ε-prediction (c = 0, d = 1), v-prediction (c = b(t), d = −a(t))

How do you choose these functions? The slides' answer: usually through some mathematical formalism. Slides 115–117 give three views:

ViewWhat the slides sayKey papers
Latent variable modelThe forward process (adding Gaussian noise) is known; learn a network to approximate the backward process and optimize a variational lower bound (same as a VAE)Sohl-Dickstein 2015, DDPM
Score functionThe score is the gradient of log p(x) with respect to x, a vector field pointing toward high density; diffusion learns the score of p_dataSong & Ermon 2019, DDPM
Stochastic differential equationsWrite the continuous noising process as an SDE; diffusion learns to approximately solve itSong et al. 2021

Slide 118 recommends Sander Dieleman's Perspectives on diffusion, adding "all his blog posts are great".

Back to the models: autoregression strikes back

Slides 119–120 pay off the previous lecture's setup. Autoregressive models are too slow on raw pixels, but work great on discrete latents: train an encoder and decoder that turn an image into an H/D×W/D grid of integers, model that sequence of discrete tokens autoregressively, then sample from the autoregressive model and pass the result to the decoder. The line of work cited is VQ-VAE, VQ-VAE-2, and Taming Transformers.

Taken together, the two lectures show that none of the four paradigms wiped out the others: GANs became a training component for the LDM decoder, VAEs became the latent space, autoregression came back on discrete latents, and diffusion is today's main engine for generating images and video.

Watch the assignment mapping. Q3 of A3 asks you to implement DDPM (DDPM.ipynb, with unet.py and gaussian_diffusion.py in the starter code), not the rectified flow at the center of the lecture. The bridge is the generalized framework above: a comment in the starter gaussian_diffusion.py reads x_t = √ᾱ_t · x₀ + √(1 − ᾱ_t) · noise, which is the VP form, and objective defaults to pred_noise (ε-prediction), with pred_x_start (x-prediction) as the other option. Reading slides 110–115 before the assignment makes the notation easier to line up. Details are in the A3 guide.

Going deeper

Official material

Related posts on this site (each is self-contained; overlap is kept)

Access limits

Under the grading in the global AI/CS course map, this course is A3: the 2026 slides, assignments, and starter code are public, plus full 2025 recordings. The gaps for this lecture: 2026 recordings are on Canvas for enrolled students only, and Gradescope autograding is not public.

References