Skip to content

CMU 10-423 L6: Generative Adversarial Networks and Probabilistic Graphical Models

Sep 30, 20261 min
TL;DRL6 is the first real generative model in 10-423's image unit. A GAN is two deterministic networks: a generator that turns Gaussian noise into an image and a discriminator that tells real from fake. They play a minimax game and take turns with mini-batch SGD updates. The deck then covers scale, watermarking and societal impact, and closes with directed graphical models, Markov models and factor graphs to set up L7's diffusion models.

🌏 中文版

Version note: This post is based on the Spring 2026 offering of CMU 10-423/623/723 Generative AI, co-taught by Aran Nayebi and Matt Gormley. The main materials are the Lecture 6 slides (a 75-page PDF) and the two readings listed on the schedule. All facts were checked against the official materials on 2026-09-30. Access level A3: slides, homework and the practice exam are public. Lecture recordings sit on CMU's Panopto and are not viewable off campus, so this post works from the slides only.

Series: previous L5: CNNs, encoder-only Transformers and ViT | next L7: An introduction to diffusion models | Series overview

L5 was about understanding images. L6 (February 2, 2026) flips the question: can a model draw an image it has never seen? This lecture's answer is the Generative Adversarial Network (GAN). The deck is titled "GANs + Probabilistic Graphical Models." The first half covers GANs. The second half is a crash course in graphical models, which looks like a detour until L7 uses exactly those diagrams to describe diffusion.

Five image generation tasks

The slides open with five tasks, each illustrated with a figure from a paper:

  • Class-conditional generation: given a label such as "sea anemone" or "goldfinch," sample a new image of that class. The slides frame it as classification in reverse: classification is p(y|x), this is p(x|y).
  • Super resolution: reconstruct a high-resolution image from a low-resolution one.
  • Image editing: inpainting (fill in specified missing pixels), colorization (add color to a grayscale image), and uncropping (fill in a missing side of an image).
  • Style transfer: keep the semantic content of a source image but render it in the style of another.
  • Text-to-image: generate an image that matches a text prompt. The examples come from Gemini 2.5 Flash Image and SDXL.

There is also a small joke: Stable Diffusion, asked to draw "a slide explaining GANs," can't do it. GPT-5 does better but skips the details. This lecture supplies the details.

A GAN is two networks

A GAN contains two deterministic neural networks:

  1. Generator G_θ: takes a random noise vector z (usually z ~ N(0, σ²I)) and outputs an image x = G_θ(z). The slides use DCGAN as the example: an "inverted CNN" whose four fractionally-strided convolution layers grow the image layer by layer, ending in three RGB channels.
  2. Discriminator D_φ: takes an image and outputs the probability that it is real, p(real | image), with real labeled 1 and fake labeled 0. The example is PatchGAN, which classifies each image patch as real or fake. According to the slides, this helps avoid blurry outputs.

During training the two play a minimax game. The generator tries to fool the discriminator; the discriminator tries to tell real from fake.

Objective and alternating updates

The discriminator sees two kinds of input: the generator's fake G_θ(z) and a real image x' drawn from the data distribution. The slides write the loss as a sum of two terms: J = log(1 − D_φ(G_θ(z))) for the fake and J' = log D_φ(x') for the real image.

The minimax objective (slide 31)
Discriminator: max_φ  Σ_i [ log D_φ(x^(i)) + log(1 − D_φ(G_θ(z^(i)))) ]
               maximize the likelihood of a binary classifier (real=1, fake=0)
               on the fixed output of the generator

Generator:     min_θ  Σ_i log(1 − D_φ(G_θ(z^(i))))
               minimize the likelihood that its fakes are classified as fake,
               according to a fixed discriminator

Since G and D are both differentiable networks, the objective is a simple differentiable function. Training alternates:

  • hold G_θ fixed and backprop through D_φ;
  • hold D_φ fixed and backprop through G_θ.

The slides compare this to block coordinate descent, except that each step takes one mini-batch SGD step instead of solving the min or max exactly. The training data is m unlabeled images.

The deck leaves an in-class question here: how do you backpropagate through G_θ when a random Gaussian is involved? The answer box is blank in the handout. Question 6 of the practice exam (GANs, 9 points) has similar conceptual questions you can use to check yourself.

A class-conditional GAN is straightforward: append a label embedding to the inputs of both the generator and the discriminator, and the GAN can generate a chosen class.

Scale, watermarking and societal impact

The "Scaling up" section starts with the long list of GAN variants in the-gan-zoo (the slide title is "GANs Everywhere!"). A computer vision timeline then places GANs (2014) among VAEs (2013), diffusion models (2015), DDPM (2020) and Stable Diffusion. The deck compares GAN and diffusion samples and uses Parti samples at different model sizes to show what scale buys.

Next come the problems generated images create, split into four efforts:

ApproachGoalWhat the slides say
WatermarkingTell whether an image was generated by a modelMost methods (GANs, VAEs, Stable Diffusion) can be augmented with a watermark
Fake-image detectionSpot fakes even without a watermark—
Model attributionIdentify which model made an image (e.g. DALL-E 2 vs. SDXL)Very successful; models leave "natural watermarks"
Image attributionFind the training images that led to a new imageExtremely challenging

The societal impact slide lists pros (new tools for artists, faster meme creation) and cons (copyright infringement and lost work for artists, a societal decrease in creativity, potential for dehumanizing content, fake news and harder fact checking, content not rooted in reality).

The second half: probabilistic graphical models

The last twenty or so slides are a crash course in graphical models, so that L7 can describe diffusion with pictures right away.

  • Directed graphical models (Bayesian networks): a directed acyclic graph with one node per variable. The joint distribution factorizes into a product of each variable's probability given its parents. The graph can come from domain knowledge, be learned from data, or be chosen for computational convenience; the conditional probabilities are usually learned. The slides show both a discrete version (probability tables) and a continuous one (Gaussian conditionals), and explain that shaded nodes are observed.
  • Markov models: the first-order Markov assumption says x_t is conditionally independent of earlier variables given x_{t−1}. The joint is p(x_1) times a chain of p(x_t | x_{t−1}). In L7 this chain becomes the diffusion noising process.
  • In-class exercise: draw a five-word RNN language model as a directed graphical model, which ties back to post 1.
  • Undirected graphical models: cliques, maximal cliques and separation are defined, but the slide says outright that these "are complicated and we don't really need them here."
  • Factor graphs: a bipartite graph of variables (circles) and factors (squares). Each factor's potential table scores assignments to its neighbors. Multiply the relevant factors to score an assignment; because the scores sum to some Z > 1, divide by Z to get probabilities.

Where this lecture shows up in homework and exams

  • HW2 (Generative Models of Images, 60 points total) has a 5-point GAN question: in a grayscale inpainting setting, you write a GAN-style objective.
  • Quiz 2 is on February 16 and covers L5–L9, GANs included.
  • Question 6 of the practice exam covers GANs for 9 points, and solutions are provided.

How to self-study it

  1. Read slides 19–37 (GAN structure and training), then explain each direction of the minimax objective out loud.
  2. Read Algorithm 1 in Goodfellow et al. 2014 (slide 36 reproduces it) and match it to "one SGD step at a time."
  3. For intuition, training tricks and common problems, read Goodfellow's NeurIPS 2016 GAN tutorial.
  4. From the graphical models section, learn the first-order Markov chain on slide 61. The next post uses it immediately.
  5. Finish with practice exam question 6 and check the solutions.

Further reading

References