Skip to content

CS231N L13: Generative Models I — What Autoregressive Models, VAEs, and GANs Each Optimize

Sep 30, 20261 min
TL;DRThe first generative-models lecture of CS231N Spring 2026 starts by pinning down the difference between a discriminative model, which learns p(y|x), and a generative model, which learns p(x). Every possible image competes for the same probability mass, so a generative model can reject unreasonable inputs. A taxonomy then splits generative models into those that can compute p(x) and those that can only sample. The lecture covers the two that compute (or approximate) it: autoregressive models factor p(x) with the chain rule into step-by-step predictions, and their weak point on raw pixels is speed; VAEs cannot compute p(x), so they maximize a lower bound, the ELBO, whose reconstruction and prior terms pull against each other. The schedule lists GANs under this lecture, but both the 2026 and 2025 slides place them at the start of the next one. This post covers them too so the three paradigms can be compared in one place.

🌏 中文版

Source years: slides and assignments are from Spring 2026; the recordings are from Spring 2025 (YouTube). The two may differ. This post follows the 2026 slides and uses the recording only as a supplement.

This is part 15 of the Reading Stanford CS231N series. The previous post is L12: Self-Supervised Learning; the next is L14: Generative Models II, Diffusion.

The CS231N lecture on May 14, 2026 was the first half of generative models. The schedule lists three topics: Variational Autoencoders, Generative Adversarial Network, and Autoregressive Models, with one suggested reading, the blog post ELBO — What & Why. The official material is the 116-page lecture_13.pdf; the matching public recording is Spring 2025 Lecture 13.

One thing to clear up first: the 2026 L13 slides actually cover only autoregressive models and VAEs. The taxonomy (slides 46–47) marks autoregressive models and VAEs as "Today" and GANs and diffusion as "Next Time", and the last slide reads "Next Time: Generative Adversarial Networks, Diffusion Models". The GAN material is on slides 8–35 of lecture_14.pdf. The 2025 slides split things the same way. Following the schedule and the series plan, this post includes GANs and marks where they really sit in the slides; the next post starts at diffusion.

Scene: a classifier can't say "this image makes no sense"

The slides open by revisiting a familiar contrast (slides 14–21):

  • Supervised learning gets (x, y) and learns a function x → y: classification, detection, segmentation, captioning
  • Unsupervised learning gets only x and learns hidden structure in the data: clustering, dimensionality reduction, density estimation

Slides 23–30 then sharpen this with three kinds of probabilistic model:

ModelLearnsWhat competes for probability
Discriminativep(y | x)The labels for one image compete; different images do not
Generativep(x)All possible images compete for the same probability mass
Conditional generativep(x | y)Each label sets off its own competition across all images

The key is the second column. A discriminative model must output a label distribution for any input; feed it an abstract painting and it can only split probability between "cat" and "dog". Because all images share a total probability of 1 in a generative model, it can give an unreasonable image tiny probability, in effect "rejecting" it. The slides point out this takes deep understanding: is a dog more likely to sit or stand? Is a three-legged dog more likely than a three-armed monkey?

Slide 36 adds a practical note: "generative models" can mean either the unconditional or conditional version, and conditional generative models are the most common in practice.

Why generative models at all? Slides 37–39 answer: modeling ambiguity. When one input y has many plausible outputs x, you want to model P(x | y). The three examples are language modeling (the slide has a model write a short rhyming poem about generative models), text-to-image, and image-to-video that predicts what happens next.

Intuition: one taxonomy splits generative models into two camps

The taxonomy on slide 46, adapted from Ian Goodfellow's 2017 GAN tutorial, is the map for these two lectures:

Generative models
├── Explicit density (model can compute P(x))
│   ├── Tractable: autoregressive models
│   └── Approximate: VAE
└── Implicit density (cannot compute p(x), but can sample)
    ├── Direct sampling: GAN
    └── Iterative procedure to approximate samples: Diffusion

Read it this way: the further down you go, the more the model gives up on writing p(x) down in exchange for better sampling. Autoregressive models compute probabilities honestly; VAEs can't, so they compute a lower bound; GANs don't compute it at all and only need to sample; diffusion reaches a sample through many small steps.

Mechanism 1: autoregressive models — p(x) as a chain of predictions

Maximum likelihood. Slide 52 sets the goal: write down an explicit function p(x) = f(x, W), then maximize the probability of the training data. Taking the log turns the product into a sum, giving a loss you can optimize with gradient descent.

Chain rule. If x is a sequence (x₁, …, x_T), the chain rule of probability factors the joint probability into a product of "each step given all previous steps" (slide 55). The slide notes: we have already seen this — it is language modeling with an RNN. Slide 56 adds that LLMs are autoregressive models too, just with a masked Transformer.

Formulas: maximum likelihood and the chain rule (slides 52, 55)

$$W^* = \arg\max_W \prod_i p(x^{(i)}) = \arg\max_W \sum_i \log f(x^{(i)}, W)$$

$$p(x) = p(x_1, x_2, \dots, x_T) = p(x_1),p(x_2 \mid x_1),p(x_3 \mid x_1, x_2)\cdots = \prod_{t=1}^{T} p(x_t \mid x_1, \dots, x_{t-1})$$

Applied to images. Slide 58 follows PixelRNN and PixelCNN: flatten the image into a sequence of 8-bit subpixel values in scanline order, treat each subpixel as a 256-way classification, and model it with an RNN or Transformer.

The weak point is cost. A 1024×1024 image is a sequence of 3 million subpixels. Slide 59 "jumps ahead" to a fix: model the image as a sequence of tiles rather than subpixels. This thread returns at the end of the next lecture (autoregression over discrete latents); see the next post.

Mechanism 2: from autoencoders to VAEs

Plain autoencoders. Slides 63–67 start with the non-variational version. The goal is to extract useful features z from inputs x without labels. How do you train without labels? Have a decoder reconstruct x from z and use the L2 distance between x̂ and x as the loss — "encoding yourself". After training, the encoder can be reused for downstream tasks.

Where it gets stuck. Slides 68–70 ask: since the decoder turns z into an image, can we invent a new z to generate images? The problem is that generating a new z is no easier than generating a new x. The fix: force all z to come from a known distribution. That is the starting point of the VAE (Kingma & Welling).

The probabilistic story (slides 75–86):

  1. Assume each training image x is generated from an unobserved latent representation z; think of z as attributes, orientation, and other factors that produce x
  2. Assume a simple prior p(z), such as a Gaussian
  3. To generate: sample z from the prior, then sample x from the conditional p(x | z)

The natural way to train is maximum likelihood. But z is unobserved, so you have to marginalize it out, and you can't integrate over all z. Bayes' rule doesn't help either, because the true posterior p(z | x) is intractable. The fix is to train a second network q(z | x) that approximates p(z | x) — the encoder.

How does a network output a distribution? Slides 89–90: the network outputs the mean (and standard deviation) of a (diagonal) Gaussian. The decoder outputs a mean for x with fixed variance, so maximizing log p(x | z) is equivalent to minimizing the L2 distance between x and the network output.

Formulas: the ELBO derivation (slides 92–102)

Start from Bayes' rule and multiply top and bottom by q(z | x):

$$\log p_\theta(x) = \log \frac{p_\theta(x \mid z),p(z)}{p_\theta(z \mid x)} = \log \frac{p_\theta(x \mid z),p(z),q_\phi(z \mid x)}{p_\theta(z \mid x),q_\phi(z \mid x)}$$

Rearrange the logs. The left side doesn't depend on z, so it can be wrapped in an expectation over z ∼ q(z | x):

$$\log p_\theta(x) = \mathbb{E}{z}[\log p\theta(x \mid z)] - D_{KL}\big(q_\phi(z \mid x),|,p(z)\big) + D_{KL}\big(q_\phi(z \mid x),|,p_\theta(z \mid x)\big)$$

The three terms are:

  • Data reconstruction: x passed through the encoder and decoder should come back
  • Prior: the encoder output should match the prior over z; closed form when both are Gaussian
  • Posterior approximation: the encoder output should match the true posterior; this term cannot be computed

KL divergence is always ≥ 0, so dropping the last term gives a lower bound:

$$\log p_\theta(x) \ge \mathbb{E}{z \sim q\phi(z \mid x)}[\log p_\theta(x \mid z)] - D_{KL}\big(q_\phi(z \mid x),|,p(z)\big)$$

The last line of that block is the VAE training objective, the variational lower bound, also called the Evidence Lower Bound (ELBO). The encoder and decoder are trained jointly to make it as large as possible. The suggested ELBO blog post opens by describing exactly this role: it turns an intractable inference problem into an optimization problem you can solve with gradient methods.

One training iteration (slide 109):

  1. Run x through the encoder to get a distribution over z
  2. Prior loss: the encoder output should be close to a unit Gaussian (zero mean, unit variance)
  3. Sample z from the encoder output using the reparameterization trick: draw ε ∼ N(0, I), then compute z = ε ⊙ Σ + μ
  4. Run z through the decoder to get the predicted data mean
  5. Reconstruction loss: the predicted mean should match x in L2

The two losses fight. Slide 110 names the VAE's central tension: the reconstruction loss wants Σ = 0 and a unique μ for every x so the decoder can reconstruct deterministically; the prior loss wants Σ = I and μ = 0 so the encoder output is always a unit Gaussian. The trained model is a compromise between the two.

Sampling and disentangling. To generate, sample z from the prior N(0, I) and run the decoder once (slide 111). Because the prior is a diagonal Gaussian, the dimensions of z are independent, so varying z₁ or z₂ alone shows different factors of variation — what the slides call "disentangling factors of variation" (slide 112).

Mechanism 3: GANs — give up on p(x), just sample (slides are in L14)

This section comes from slides 8–35 of lecture_14.pdf; the recording is Spring 2025 Lecture 14.

Slide 10 lines up the first two models: autoregressive models directly maximize the likelihood of the training data; VAEs introduce a latent z and maximize a lower bound. GANs give up on modeling p(x) but let us draw samples from it.

Setup (slide 15): data x comes from p_data. Introduce a simple prior p(z), sample z, pass it to a generator G to get x = G(z), which follows the generator distribution p_G. We want p_G = p_data. To get there, train a discriminator D to tell real from fake, and train the generator by fooling the discriminator.

A minimax game (slides 19–22): D(x) is the probability that x is real. The discriminator wants real data classified as 1 and fake data as 0; the generator wants fake data classified as 1. The two are trained with alternating gradient updates. The slides flag an important point: we are not minimizing any overall loss, so there are no training curves to look at.

Vanishing gradients early on (slide 25): at the start the generator is bad and the discriminator easily tells real from fake, so D(G(z)) is near 0 and the generator's gradients are near 0 too. The fix is to have the generator minimize −log D(G(z)) instead, which gives strong gradients early.

Formulas: the GAN objective and its optimum (slides 19, 29)

$$\min_G \max_D; \mathbb{E}{x \sim p{data}}[\log D(x)] + \mathbb{E}_{z \sim p(z)}[\log(1 - D(G(z)))]$$

Alternating updates: D ← D + α_D ∂V/∂D, G ← G − α_G ∂V/∂G.

For any fixed p_G, the optimal inner discriminator is:

$$D^*G(x) = \frac{p{data}(x)}{p_{data}(x) + p_G(x)}$$

Plugging it back in, the outer objective is minimized at p_G = p_data (the slides omit the proof).

That optimum shows the objective itself is sound, but the slides list two caveats: fixed-capacity networks may not be able to represent the optimal D and G, and the result says nothing about convergence with finite data.

Architectures (slides 31–34):

  • DC-GAN (Radford et al., ICLR 2016): G and D are both usually CNNs; it was the first GAN architecture that worked on non-toy data. The slides also note that GANs fell out of favor before ViT became popular
  • StyleGAN: a more complex generator that injects noise via adaptive normalization, predicting at each layer a scale w and shift b of the same shape as x
  • Latent-space interpolation: a GAN's latent space is smooth; interpolating linearly between z₀ and z₁ gives images that transition smoothly

GAN summary (slide 35): pros are a simple formulation and very good image quality; cons are no loss curve to look at, unstable training, and difficulty scaling to big models and data. The slides' verdict: GANs were the go-to generative models from about 2016 to 2021.

Back to the models: what each paradigm optimizes

AutoregressiveVAEGAN
ObjectiveExact log p(x) (chain rule)ELBO, a lower bound on p(x)Minimax game with a discriminator
Computes p(x)?YesOnly approximatelyNo
SamplingGenerate step by stepSample z from prior, decode onceSample z from prior, generate once
Pain point named in slidesToo slow on raw pixelsReconstruction and prior losses fightNo loss curve, unstable, hard to scale

The table also previews the next lecture: diffusion is the last box in the taxonomy, "iterative procedure to approximate samples", and modern latent diffusion brings back both VAEs and GANs as components. Autoregression returns on discrete latents.

On the assignment side, A3 has no VAE or GAN problem; its generative-model exercise is the DDPM in Q3, covered in the A3 guide.

Going deeper

Official material

Related posts on this site (each is self-contained; overlap is kept)

Access limits

Under the grading in the global AI/CS course map, this course is A3: the 2026 slides, assignments, and course notes are public, plus full 2025 recordings. The gaps for this lecture: 2026 recordings are on Canvas for enrolled students only, and the midterm (May 12, two days before this lecture) is not public.

References