Skip to content

CMU 10-423 HW2: Implementing DDPM from Scratch on AFHQ Cats — Structure, Files to Edit, and Compute

Sep 30, 20261 min
TL;DRHW2 in the Spring 2026 CMU 10-423 is worth 60 points. The written part covers CNNs (8), encoder-only Transformers (4), GANs (5), VAEs (6), and diffusion models (14). The programming part (21) has you implement DDPM from scratch on AFHQ cat images: fill in the TODOs in diffusion.py and unet.py, then submit loss curves, FID curves, and forward/reverse diffusion figures from W&B. The longest experiment trains for 10,000 steps, which the handout estimates at about 2 hours on a Colab T4.

🌏 中文版

This post is based on the Spring 2026 offering of CMU 10-423/623/723 Generative AI. It is part 9 of the Reading CMU 10-423 series and closes the image unit. The previous four posts covered CNNs and ViT, GANs, diffusion models, and variational inference and VAEs. This one looks at how the homework tests all of them, then asks you to write a DDPM that draws cats.

Official materials used: hw2.zip from the Coursework page (32 MB: the 27-page hw2.pdf, starter code, data, and a LaTeX template), the HW2 read-only Overleaf template, the February 13 HW2 recitation slides (a public PowerPoint file on Google Drive), and the homework rules in the syllabus. I checked every fact against these materials on 2026-09-30. This post covers the structure and setup only. It contains no solutions.

The basics

ItemDetails
NameHomework 2: Generative Models of Images (listed as "Image Generation" on the Coursework page)
ScopeL5–L8 (the schedule says "HW2 out (L5-L8)" on February 9)
Released2026-02-09
Due2026-02-21 at 11:59pm (Slot A, per the Reminders slide in Lecture 10)
Slot BTentatively March 1 on the schedule; feedback tentatively February 26
SubmissionGradescope: one written PDF; code is diffusion.py only
Total60 points

Point breakdown (hw2.pdf, page 2):

SectionPoints
LaTeX Template Alignment0
Convolutional Neural Networks8
Encoder-only Transformers4
Generative Adversarial Network (GAN)5
Variational Autoencoders6
Understanding Diffusion Models14
Programming: Diffusion Models21
Code Upload0
Collaboration Questions2

Written answers must follow the template. A misaligned submission loses 2% of the assignment's points, and hw2.pdf explains why: grading uses an AI-assisted grader. The syllabus's Slot A/B rules also apply. Slot A takes pure human work only. Three days after feedback, you can fix the questions you got wrong in Slot B with AI help, and each question keeps the higher of the two scores. The full rules are in the series overview.

What the written questions test

The five written sections map almost one-to-one onto L5–L8, in the same order.

CNNs (8 points). You get a convolution definition with a single output channel and work out how many columns and rows of "hallucinated pixels" are needed on each side, using only floor functions. Then you rewrite the equations so the input is explicitly zero-padded and only valid positions are indexed. The last two parts are conceptual questions about U-Net: the role of skip connections, and how feature-map resolution changes through the encoder and decoder. These set up the programming part, because the DDPM denoiser is a U-Net.

Encoder-only Transformers (4 points). Draw the directed graphical model for the sequence distribution a decoder-only Transformer defines. Then compare it with an encoder-only model plus a per-word part-of-speech head: what kinds of sequence distributions each can model, and what each is suited for. This ties back to the BERT part of L5.

GANs (5 points). A character in the problem, Lora the Llama, wants to inpaint grayscale photos of fellow llamas. You fill in three blanks in a training pseudocode: a model output that replaces only the masked pixels, a squared error loss over the masked region only, and a GAN-style objective. The last part asks for one drawback of training with squared error alone.

VAEs (6 points). The first part gives a one-dimensional linear-Gaussian encoder and decoder and asks which parameters receive gradients from the reconstruction term under two computation graphs: black-box sampling versus the reparameterization trick. The second part applies a VAE to video and asks for the exact loss on one batch (an MSE reconstruction term plus a β-weighted KL term). The third is a multiple-choice question on β-VAE: as β grows from 0, what happens to the latent representation and to reconstruction quality?

Understanding Diffusion Models (14 points). The heaviest written section, in two parts:

  • ELBO Surgery (6.1–6.2). Question 6.1 (5 points) asks you to start from the ELBO and show it splits into a reconstruction term, a sum of L_t terms, and a constant that doesn't depend on θ. The problem gives six hints, adds a note that "proofs are scary," and expects an answer that is mostly equations. Question 6.2 asks what L_t does to the reverse process and when it is maximized.
  • Image Diffusion (6.3–6.5). Given the per-step forward distribution, describe how to sample step by step up to step τ and state the time complexity. Then use the reparameterization trick to derive the closed form that jumps straight from x₀ to x_t. Finally, describe sampling with that closed form and its complexity.

The formula you derive in 6.4 is exactly what q_sample implements in the programming part, so it makes sense to finish section 6 before touching the code. The ELBO background is in part 8.

The programming part: DDPM from scratch

Data

hw2.pdf describes the dataset as Animal Faces-HQ (AFHQ), "15,000 images at 36 × 36 resolution," in three domains: cat, dog, and wildlife. The assignment uses only the cat images to keep compute down. The handout also warns up front that because this is a subset of the original dataset, generated images may not look as good as you'd hope. I unzipped the archive and counted: data/train/ has 5,153 cat, 4,739 dog, and 4,738 wild images, data/val/ has 500 of each, and every image is indeed 36 × 36. The model's default image_size is 32.

The data ships inside the zip, so there's nothing extra to download. The Kaggle notebook also points to a pre-uploaded Kaggle dataset, which still loaded on 2026-09-30.

Files and what to change

FilePurpose
diffusion.pyThe Diffusion class: forward process, reverse process, and schedule. You edit this.
unet.pyThe U-Net denoiser. You edit this (only forward).
trainer.pyTraining loop, sampling, FID, W&B logging
utils.pyThe train_diffusion and visualize_diffusion entry points plus W&B setup
main.pyCommand-line entry point for local or AWS runs
run_in_colab.ipynb / run_in_kaggle.ipynbCloud notebooks
test_diffusion.py, data.ptUnit tests and their fixed test tensors
requirements.txtDependencies

hw2.pdf says diffusion.py and unet.py are the only files you need to modify, and every spot is marked TODO. Opening the files:

  • diffusion.py has 7 TODOs: the precomputed coefficients in __init__, plus six functions: q_sample, p_losses, forward, p_sample, p_sample_loop, and sample. The recitation slides group them into training (forward, p_losses, q_sample) and sampling (sample, p_sample_loop, p_sample).
  • The noise schedule is already written. cosine_schedule implements the cosine schedule from Nichol & Dhariwal (2021) with s = 0.008.
  • The TODOs in unet.py are all in Unet.forward. On the downsampling path, each level runs two residual blocks and attention and saves skip features. Then comes the bottleneck. On the upsampling path, the skip features are concatenated back in. The building blocks (ResnetBlock, LinearAttention, SinusoidalPosEmb, and so on) are provided.

Training details fixed by hw2.pdf: p_losses compares the true and predicted noise with an L1 loss. The sampling algorithm first estimates x̂₀ from the predicted noise, clamps it to [−1, 1], then takes one step using the posterior mean and variance. trainer.py uses Adam.

One thing to watch: the Code Upload question and the test file both say to upload only diffusion.py, yet unet.py also has TODOs and the eighth test checks the U-Net's forward. How Gradescope actually handles unet.py isn't visible from outside.

Default hyperparameters

ParameterDefault
time_steps (T)50
batch_size32
image_size32
unet_dim16
unet_dim_mults[1, 2, 4, 8]
learning_rate1e-3
data_classcat
dataloader_workers16

hw2.pdf says the values in its Table 4 don't need to change for this homework. Between experiments you only adjust train_steps, save_and_sample_every, and fid.

Unit tests

test_diffusion.py has 8 tests, all forced onto the CPU (a comment says the GPU may give slightly different outputs). T01 (checks submitted files) and T02 (noise_like) have weight 0. T03–T08 are worth 1 point each and test p_sample, p_losses, p_sample_loop, sample, q_sample, and the U-Net forward.

On 2026-09-30 I ran the unmodified starter code on macOS with Python 3.12 and PyTorch 2.14.0 (CPU). All 8 tests finished in about 5 seconds: T01 and T02 passed, 4 tests failed, and 2 errored, which is what you'd expect with the TODOs still empty. Note that diffusion.py runs import wandb at the top, so you need wandb installed even to run the tests. I installed only torch, einops, tqdm, and wandb, without the version pins in requirements.txt.

The five experiment questions

All 21 programming points come from experiment results, and the figures come from W&B:

QuestionTaskSuggested parametersHandout estimate (Colab T4)Points
7.1Training loss over 1,000 steps; the model should produce blurry cats by thentrain_steps=1000, save_and_sample_every=100, fid=False5–10 minutes4
7.2FID every 100 steps, plotted over 1,000 stepssame, with fid=True15–60 minutes4
7.3Train the full 10,000 steps and show the last sample batchtrain_steps=10000, save_and_sample_every=1000, fid=Falseabout 2 hours5
7.4Using the 10,000-step model, show the forward process at 0%, 25%, 50%, 75%, and 99%call visualize_diffusion—4
7.5Feed the last forward-process noise back in and show the reverse process at the same five pointssame—4

FID is computed with compute_fid from clean-fid. hw2.pdf describes the comparison as being against resized training images. trainer.py, however, builds the reference folder by replacing train with val in data_path, and only fills it with resized training cats if that folder is empty. The zip already ships a non-empty data/val/. I didn't run FID, so I haven't confirmed which images the default paths actually compare against.

Environment and compute: can you still run it today?

hw2.pdf lists three ways to run it: main.py locally or on AWS, or the Colab or Kaggle notebook in the cloud. For step-by-step GPU setup it points back to the HW1 handout. Both notebooks work the same way: install requirements.txt, paste your entire diffusion.py into the designated cell, fill in train_steps, save_and_sample_every, and fid, then call train_diffusion and visualize_diffusion.

Budget your compute generously:

  • The recitation slides call this the homework with the most training time in the course. The required experiments take 2–3 hours on a Colab T4, and you'll likely rerun them while debugging. The advice: start early, and test your code with a few steps first.
  • One slide in the same deck says "Colab Pro is FREE for students." That's addressed to enrolled students. I didn't check whether it applies to readers outside CMU.
  • requirements.txt requires Python 3.9 or newer but below 3.13, and pins clean-fid 0.1.35, einops 0.8.1, pillow 11.3.0, tqdm 4.67.1, and wandb 0.21.3. torch and torchvision are unpinned.
  • utils.py calls wandb.login() before both training and visualization, so you need a Weights & Biases account. Runs log to a project named DDPM_AFHQ.

Like HW1, HW2 needs no external weights, and the data is in the zip. The extra hurdles are a W&B account and more GPU time.

What the recitation slides add

The February 13 recitation deck has 75 slides in three parts:

  • Starter code walkthrough. It explains each file and splits the implementation into four pieces: the U-Net's forward, the noise schedule (already done), the training algorithm, and the sampling algorithm. The last two are the six functions in diffusion.py.
  • Written-section review. It explains diffusion with an add-noise/remove-noise analogy and stresses that "denoising is not image recovery… it's image generation!" It introduces FID: features from Inception-v3, lower is closer to the target distribution. For the ELBO, it uses the analogy of a crime-scene reporter who must stay accurate but also wants clicks, to explain the reconstruction and KL terms. It then derives the ELBO step by step with Jensen's inequality and ends with the reparameterization trick.
  • Helpful functions. Why use noise_like instead of calling torch.randn directly (the slide's answer: it keeps randomness consistent across submissions), how gather_timestep_coeff pulls the needed timesteps out of precomputed coefficient tables, and quick quizzes on torch.cumprod, torch.clamp, and torch.full. The torch.clamp quiz deliberately shows that passing a numpy array raises a TypeError.

The deck is a PowerPoint file on Google Drive rather than a native Google Slides deck. On 2026-09-30 it could be exported and downloaded directly.

How this homework connects to the tests

The schedule puts "Programming Test HW1/HW2" in class on February 25, covering the programming parts of HW1 and HW2, closed-book. The test itself isn't available outside CMU, but it shows the course expects you to understand your own q_sample and p_sample, not just make the tests pass. In the practice exam, the questions in HW2's scope are ViT (14 points), GANs (9), VAEs (8), and Diffusion Models (8).

What to do: tonight, download hw2.zip, install torch, einops, tqdm, and wandb, and run python test_diffusion.py. Confirm that T01 and T02 pass and the rest fail. Then do written question 6.4 first to derive the form of q(x_t | x₀), open diffusion.py, and fill in __init__ and q_sample until T07 turns green.

What this post can and can't confirm

Confirmed: the PDF, starter code, data, and tests in hw2.zip; the recitation slides; the syllabus's Slot A/B rules; the dates on the schedule; and the results of running the unit tests locally. Not confirmed: Gradescope's autograding and manual grading criteria, how unet.py is graded on Gradescope, whether the tentative Slot B date on the schedule (March 1) held that semester, which reference images FID actually uses, real training time on a Colab T4 (I didn't run on a GPU), and the official solutions.

Further reading: to revisit diffusion and flow matching through ODEs and SDEs, see Reading MIT 6.S184. Its Lab 3 is another from-scratch generative-model assignment worth comparing.

Series navigation: previous L8–L9: variational inference, VAEs, and the diffusion ELBO | next L10–L11: parameter-efficient fine-tuning and in-context learning | series overview

References