The last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).
CMU 10-423 Lecture 5 opens the image unit with three models built for understanding. CNNs treat the convolution kernel as parameters to learn. Encoder-only Transformers drop the causal mask so every token sees both sides and train with a masked LM objective, which means they are not generative language models. ViT is nearly BERT with image patches such as 16×16 pixels as input. The slides use a figure from the ViT paper to explain why Transformers reached vision four years after NLP: on small datasets ViT loses to large CNNs, and it only pulls ahead with enough data.
CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.
Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.
L7 splits a diffusion model into two Markov chains. A fixed forward process gradually turns an image into Gaussian noise, and a learned reverse process removes the noise step by step. The exact reverse process is intractable, but the posterior given the original image x₀ is a closed-form Gaussian, so it can serve as the learning target. The slides compare three parameterizations. The best in practice has a U-Net predict the noise ε that was added, and the training loop is eight lines long.
The last two lectures of the Scaling Up unit in CMU 10-423 Spring 2026. L17 starts from the claim that communication between GPUs is the main bottleneck, then walks through data parallelism, Megatron-style tensor parallelism, 1F1B pipeline parallelism, ZeRO optimizer parallelism, TeraPipe token parallelism, and expert parallelism. Its conclusion: data parallelism is still king, and the rest exist to push more data through it. L18 covers two things. FlashAttention combines tiling, online softmax, and recomputation to cut HBM traffic without changing the result. On the decoding side, PagedAttention manages KV-cache memory and speculative decoding reduces calls to the large model.
Beyond its four homework assignments, CMU 10-423 checks learning four ways: 6 in-class quizzes, 2 programming tests, one comprehensive exam, and a three-person final project worth 25%. 10-623/723 students also do HW623, a paper presentation. Outside CMU you can get the practice exam with solutions (13 sections, 167 points), the HW623 handout with its 33-paper list, and the 12-page project handout. This post lays out their structure and rules and gives a self-check routine that works without peeking at the answers.
L6 is the first real generative model in 10-423's image unit. A GAN is two deterministic networks: a generator that turns Gaussian noise into an image and a discriminator that tells real from fake. They play a minimax game and take turns with mini-batch SGD updates. The deck then covers scale, watermarking and societal impact, and closes with directed graphical models, Markov models and factor graphs to set up L7's diffusion models.
CMU 10-423/623/723 is the generative AI course co-taught by Matt Gormley and Aran Nayebi. The Spring 2026 edition covers text models, image generation, adapting foundation models, multimodal models, scaling, and advanced topics in 26 lectures. The slides, the HW1–HW4 handouts and starter code, a practice exam with solutions, and the project handout are all public, which earns an A3 rating. What you cannot get: the Panopto recordings, the HW0 handout, the HW3/HW4 recitation slides, the quizzes, and Gradescope grading. The homework policy is worth a look on its own: every assignment is submitted twice, first as human-only work, then with AI allowed.
HW1 in CMU 10-423 Spring 2026 is worth 62 points. The written part covers RNN LMs (7), Transformer LMs (19), and sliding window attention (11). The programming part (22) has you implement RoPE and GQA in Karpathy's minGPT, train a character-level model on the complete works of Shakespeare, and plot loss and attention time. You upload only model.py; the handout ships five unit tests, and the official estimates put all experiments at about 40 minutes on a Colab T4.
HW2 in the Spring 2026 CMU 10-423 is worth 60 points. The written part covers CNNs (8), encoder-only Transformers (4), GANs (5), VAEs (6), and diffusion models (14). The programming part (21) has you implement DDPM from scratch on AFHQ cat images: fill in the TODOs in diffusion.py and unet.py, then submit loss curves, FID curves, and forward/reverse diffusion figures from W&B. The longest experiment trains for 10,000 steps, which the handout estimates at about 2 hours on a Colab T4.
HW3 in CMU 10-423 Spring 2026 is worth 66 points and was due 2026-03-12 (Slot A). The written part covers in-context learning (14 points), parameter-efficient fine-tuning (10), and the DPO derivation (15). The programming part (25) has you write LoRALinear from scratch, wire it into GPT-2's attention, and instruction-tune the model for sentiment classification on Rotten Tomatoes reviews. Every experiment uses gpt2-medium; the handout estimates 25–30 minutes per training run on a Colab T4, and you need a WandB account.
HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.
A pretrained LLM continues text; it doesn't hold a conversation. The second half of CMU 10-423 L11 covers instruction fine-tuning, which turns the model into a chat assistant using data such as InstructGPT's 13k examples, Dolly's 15k, or Flan. Then come InstructGPT's three RLHF steps: humans rank responses, a reward model is trained, and PPO fine-tunes the policy. The first half of L12 adds the intuition behind REINFORCE and PPO, lists five drawbacks of PPO-based RLHF, and derives DPO: start from the Bradley–Terry model, replace the reward model with the policy's own log-probability ratios, and fine-tune directly on preference data.
Lectures 19 and 21 of CMU 10-423 (Spring 2026) tackle the same problem: once a sequence gets long, the memory of standard attention and the KV cache stop fitting. L19 offers two routes: approximate attention with sparse, sliding window, or dilated patterns, or keep full attention and split the computation across GPUs with the Blockwise Parallel Transformer and Ring Attention. L21 offers a third: replace attention with state space models (S4, Mamba) that keep only a fixed-size hidden state, or interleave attention with linear attention layers in hybrid models (Jamba, Nemotron-H, Qwen3-Next).
The first half of CMU 10-423 Lecture 4 separates pre-training, mid-training, and post-training. The second half picks three components that nearly every modern LLM uses. RoPE turns position into a rotation of queries and keys, so attention scores depend only on the relative distance between two tokens. GQA lets several query heads share one key/value head to save memory and compute. Sliding window attention changes the mask so each token sees only a fixed number of tokens to its left. All three show up in HW1.
With a small labeled dataset and an LLM with billions of parameters, CMU 10-423 offers two routes: supervised fine-tuning, or putting the examples in the prompt for in-context learning. L10 first notes that the 2023 consensus was that fine-tuning usually wins, then covers four ways to tune only a few parameters: the top layers only, adapters, prefix tuning, and LoRA. The first half of L11 returns to in-context learning: how sensitive it is to example order and label balance, how to pick a prompt, and what chain-of-thought is. HW3's written questions and its LoRA programming task both draw on these two lectures.
Lecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.
Lecture 22 of CMU 10-423 (Spring 2026) runs five generative AI risks through the same four questions (what is it, who does it affect, why does it happen, how do we fix it): copyright infringement, adversarial attacks, hallucination, bias and discrimination, and environmental impact, and each section ends on a concrete example of why fixing it is hard. The second deck of Lecture 26 goes a level up: Aran Nayebi uses an agreement framework to show that the cost of alignment grows with the number of tasks, agents, and state space size, so objectives must be compressed and critical states prioritized, and he proposes a lexicographic utility that puts deference and the off switch first for provable corrigibility. Data contamination, listed in the course description, does not appear in either deck.
Lecture 1 of CMU 10-423 boils generative AI down to one line: it is probabilistic modeling, and text generation means estimating p(next word | all previous words). The slides go from n-grams, which you learn by counting, to RNNs, which squeeze the previous words into a fixed-length vector. In between comes module-based autodiff: if every module can run forward and backward, gradients flow back through the computation graph automatically, and that is how PyTorch works. The HW0 handout on Google Drive returns 401; only the recitation Colab is public, covering PyTorch, LSTMs, Weights & Biases and einops.
The first two lectures of the Scaling Up unit in CMU 10-423 Spring 2026 answer two questions. The second half of L15 covers scaling laws: Kaplan 2020 says 8x more parameters needs only about 5x more data, Chinchilla says scale both equally, and the Phi models and data-filtering scaling laws add data quality as a third axis. L16 covers MoE: feed-forward layers hold most of GPT-3's parameters, so split them into experts and send each token through only the top k. Memory follows total parameters, compute follows active parameters, and the price is load balancing and training stability. No programming homework covers this half of the course; quizzes, practice exam question 13, and the final project do.
CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.
Lectures 2 and 3 of CMU 10-423 swap the RNN for attention. Lecture 2 first explains why RNNs fall short: they forget, they compute one step at a time, and their gradients can still explode. It then assembles a Transformer language model piece by piece: scaled dot-product attention, multi-head attention, layer norm, residual connections, position embeddings, and the causal mask. Lecture 3 covers training. There is no closed-form answer like n-gram counting, so you do maximum likelihood with autodiff and mini-batch SGD. It then covers padding, the KV cache and three kinds of tokenizer, and ends with greedy decoding and ancestral sampling to show how text is generated one token at a time.
VAEs and diffusion models get stuck in the same place: log p_θ(x) requires integrating over latent variables, which is intractable. L8–L9 answer with variational inference. Pick a tractable q to approximate the true posterior, and swap 'minimize the KL' for 'maximize the ELBO,' which is a lower bound on log p(x). Add Monte Carlo estimation and the reparameterization trick, and a VAE trains with one forward and one backward pass. Unpack DDPM's ELBO and every term asks the learned reverse step to match the closed-form q(x_{t−1} | x_t, x₀).