Skip to content
All tags

#cmu-10423

24 posts

CMU 10-423 L24–L26: Audio, Video Generation, and Interactive World Models — Taking Generative Models from Images to Sound, Time, and Worlds You Can Act In

The last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).

CMU 10-423 L5: CNNs, Encoder-only Transformers, and ViT — Why a Generative AI Course Starts Images with Understanding

CMU 10-423 Lecture 5 opens the image unit with three models built for understanding. CNNs treat the convolution kernel as parameters to learn. Encoder-only Transformers drop the causal mask so every token sees both sides and train with a masked LM objective, which means they are not generative language models. ViT is nearly BERT with image patches such as 16×16 pixels as input. The slides use a figure from the ViT paper to explain why Transformers reached vision four years after NLP: on small datasets ViT loses to large CNNs, and it only pulls ahead with enough data.

CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop

CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.

CMU 10-423 L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and Q-Former

Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.

CMU 10-423 L7: Diffusion Models, from Adding Noise to Learning to Remove It

L7 splits a diffusion model into two Markov chains. A fixed forward process gradually turns an image into Gaussian noise, and a learned reverse process removes the noise step by step. The exact reverse process is intractable, but the posterior given the original image x₀ is a closed-form Gaussian, so it can serve as the learning target. The slides compare three parameterizations. The best in practice has a U-Net predict the noise ε that was added, and the training loop is eight lines long.

CMU 10-423 L17–L18: Distributed Training, FlashAttention, and Efficient Decoding — Where to Start When One GPU Can't Hold the Model and Inference Is Too Slow

The last two lectures of the Scaling Up unit in CMU 10-423 Spring 2026. L17 starts from the claim that communication between GPUs is the main bottleneck, then walks through data parallelism, Megatron-style tensor parallelism, 1F1B pipeline parallelism, ZeRO optimizer parallelism, TeraPipe token parallelism, and expert parallelism. Its conclusion: data parallelism is still king, and the rest exist to push more data through it. L18 covers two things. FlashAttention combines tiling, online softmax, and recomputation to cut HBM traffic without changing the result. On the decoding side, PagedAttention manages KV-cache memory and speculative decoding reduces calls to the large model.

CMU 10-423 Wrap-up: The Practice Exam, the HW623 Paper Presentation, and the Final Project — How the Course Checks Learning, and How to Check Yourself

Beyond its four homework assignments, CMU 10-423 checks learning four ways: 6 in-class quizzes, 2 programming tests, one comprehensive exam, and a three-person final project worth 25%. 10-623/723 students also do HW623, a paper presentation. Outside CMU you can get the practice exam with solutions (13 sections, 167 points), the HW623 handout with its 33-paper list, and the 12-page project handout. This post lays out their structure and rules and gives a self-check routine that works without peeking at the answers.

CMU 10-423 L6: Generative Adversarial Networks and Probabilistic Graphical Models

L6 is the first real generative model in 10-423's image unit. A GAN is two deterministic networks: a generator that turns Gaussian noise into an image and a discriminator that tells real from fake. They play a minimax game and take turns with mini-batch SGD updates. The deck then covers scale, watermarking and societal impact, and closes with directed graphical models, Markov models and factor graphs to set up L7's diffusion models.

Reading CMU 10-423 Generative AI: Series Overview — All 26 Lecture Decks and Four Homeworks Are Public, the Videos Stay Behind Panopto

CMU 10-423/623/723 is the generative AI course co-taught by Matt Gormley and Aran Nayebi. The Spring 2026 edition covers text models, image generation, adapting foundation models, multimodal models, scaling, and advanced topics in 26 lectures. The slides, the HW1–HW4 handouts and starter code, a practice exam with solutions, and the project handout are all public, which earns an A3 rating. What you cannot get: the Panopto recordings, the HW0 handout, the HW3/HW4 recitation slides, the quizzes, and Gradescope grading. The homework policy is worth a look on its own: every assignment is submitted twice, first as human-only work, then with AI allowed.

CMU 10-423 HW1: Adding RoPE and GQA to minGPT — Structure, Files to Edit, and Compute

HW1 in CMU 10-423 Spring 2026 is worth 62 points. The written part covers RNN LMs (7), Transformer LMs (19), and sliding window attention (11). The programming part (22) has you implement RoPE and GQA in Karpathy's minGPT, train a character-level model on the complete works of Shakespeare, and plot loss and attention time. You upload only model.py; the handout ships five unit tests, and the official estimates put all experiments at about 40 minutes on a Colab T4.

CMU 10-423 HW2: Implementing DDPM from Scratch on AFHQ Cats — Structure, Files to Edit, and Compute

HW2 in the Spring 2026 CMU 10-423 is worth 60 points. The written part covers CNNs (8), encoder-only Transformers (4), GANs (5), VAEs (6), and diffusion models (14). The programming part (21) has you implement DDPM from scratch on AFHQ cat images: fill in the TODOs in diffusion.py and unet.py, then submit loss curves, FID curves, and forward/reverse diffusion figures from W&B. The longest experiment trains for 10,000 steps, which the handout estimates at about 2 hours on a Colab T4.

CMU 10-423 HW3: Fine-Tuning GPT-2 with LoRA — Written Questions, Files to Edit, and Compute

HW3 in CMU 10-423 Spring 2026 is worth 66 points and was due 2026-03-12 (Slot A). The written part covers in-context learning (14 points), parameter-efficient fine-tuning (10), and the DPO derivation (15). The programming part (25) has you write LoRALinear from scratch, wire it into GPT-2's attention, and instruction-tune the model for sentiment classification on Rotten Tomatoes reviews. Every experiment uses gpt2-medium; the handout estimates 25–30 minutes per training run on a Colab T4, and you need a WandB account.

CMU 10-423 HW4: Text-to-Image with a Q-Former Between a Frozen GPT-2 and a Frozen DiT — Structure, Files to Edit, and Compute

HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.

CMU 10-423 L11–L12: Instruction Tuning, RLHF, and DPO — Make the Model Follow Instructions, Then Drop the RL

A pretrained LLM continues text; it doesn't hold a conversation. The second half of CMU 10-423 L11 covers instruction fine-tuning, which turns the model into a chat assistant using data such as InstructGPT's 13k examples, Dolly's 15k, or Flan. Then come InstructGPT's three RLHF steps: humans rank responses, a reward model is trained, and PPO fine-tunes the policy. The first half of L12 adds the intuition behind REINFORCE and PPO, lists five drawbacks of PPO-based RLHF, and derives DPO: start from the Bradley–Terry model, replace the reward model with the policy's own log-probability ratios, and fine-tune directly on preference data.

CMU 10-423 L19 + L21: Long Context and State Space / Hybrid Models — Three Ways Out When Attention Cost Grows Quadratically

Lectures 19 and 21 of CMU 10-423 (Spring 2026) tackle the same problem: once a sequence gets long, the memory of standard attention and the KV cache stop fitting. L19 offers two routes: approximate attention with sparse, sliding window, or dilated patterns, or keep full attention and split the computation across GPUs with the Blockwise Parallel Transformer and Ring Attention. L21 offers a third: replace attention with state space models (S4, Mamba) that keep only a fixed-size hidden state, or interleave attention with linear attention layers in hybrid models (Jamba, Nemotron-H, Qwen3-Next).

CMU 10-423 L4: Pre-training, Fine-tuning, and the Modern Transformer — What RoPE, GQA, and Sliding Windows Each Fix

The first half of CMU 10-423 Lecture 4 separates pre-training, mid-training, and post-training. The second half picks three components that nearly every modern LLM uses. RoPE turns position into a rotation of queries and keys, so attention scores depend only on the relative distance between two tokens. GQA lets several query heads share one key/value head to save memory and compute. Sliding window attention changes the mask so each token sees only a fixed number of tokens to its left. All three show up in HW1.

CMU 10-423 L10–L11: Parameter-Efficient Fine-Tuning and In-Context Learning — Change a Few Weights, or Just the Input?

With a small labeled dataset and an LLM with billions of parameters, CMU 10-423 offers two routes: supervised fine-tuning, or putting the examples in the prompt for in-context learning. L10 first notes that the 2023 consensus was that fine-tuning usually wins, then covers four ways to tune only a few parameters: the top layers only, adapters, prefix tuning, and LoRA. The first half of L11 returns to in-context learning: how sensitive it is to example order and label balance, how to pick a prompt, and what chain-of-thought is. HW3's written questions and its LoRA programming task both draw on these two lectures.

CMU 10-423 L20: Reasoning Models — From Chain-of-Thought to o1, DeepSeek-R1, and GRPO, Plus a Look at Mechanistic Interpretability

Lecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.

CMU 10-423 L22 + L26: Practical Risks and the Science of Alignment — Copyright, Jailbreaks, Hallucination, Bias, Carbon, and Why Alignment Has Theoretical Limits

Lecture 22 of CMU 10-423 (Spring 2026) runs five generative AI risks through the same four questions (what is it, who does it affect, why does it happen, how do we fix it): copyright infringement, adversarial attacks, hallucination, bias and discrimination, and environmental impact, and each section ends on a concrete example of why fixing it is hard. The second deck of Lecture 26 goes a level up: Aran Nayebi uses an agreement framework to show that the cost of alignment grows with the number of tasks, agents, and state space size, so objectives must be compressed and critical states prioritized, and he proposes a lexicographic utility that puts deference and the off switch first for provable corrigibility. Data contamination, listed in the course description, does not appear in either deck.

CMU 10-423 L1: RNN Language Models and Autodiff — Generative AI Starts with Predicting the Next Word (with HW0)

Lecture 1 of CMU 10-423 boils generative AI down to one line: it is probabilistic modeling, and text generation means estimating p(next word | all previous words). The slides go from n-grams, which you learn by counting, to RNNs, which squeeze the previous words into a fixed-length vector. In between comes module-based autodiff: if every module can run forward and backward, gradients flow back through the computation graph automatically, and that is how PyTorch works. The HW0 handout on Google Drive returns 401; only the recitation Colab is public, covering PyTorch, LSTMs, Weights & Biases and einops.

CMU 10-423 L15–L16: Scaling Laws and Mixture of Experts — How Big Should the Model Be, and How Do You Compute Only Part of It?

The first two lectures of the Scaling Up unit in CMU 10-423 Spring 2026 answer two questions. The second half of L15 covers scaling laws: Kaplan 2020 says 8x more parameters needs only about 5x more data, Chinchilla says scale both equally, and the Phi models and data-filtering scaling laws add data quality as a third axis. L16 covers MoE: feed-forward layers hold most of GPT-3's parameters, so split them into experts and send each token through only the top k. Memory follows total parameters, compute follows active parameters, and the price is load balancing and training stability. No programming homework covers this half of the course; quizzes, practice exam question 13, and the final project do.

CMU 10-423 L12–L13: Text-to-Image, Latent Diffusion, and Vision-Language Models

CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.

CMU 10-423 L2–L3: Transformer Language Models, LLM Training and Decoding — From Forgetful RNNs to the KV Cache

Lectures 2 and 3 of CMU 10-423 swap the RNN for attention. Lecture 2 first explains why RNNs fall short: they forget, they compute one step at a time, and their gradients can still explode. It then assembles a Transformer language model piece by piece: scaled dot-product attention, multi-head attention, layer norm, residual connections, position embeddings, and the causal mask. Lecture 3 covers training. There is no closed-form answer like n-gram counting, so you do maximum likelihood with autodiff and mini-batch SGD. It then covers padding, the KV cache and three kinds of tokenizer, and ends with greedy decoding and ancestral sampling to show how text is generated one token at a time.

CMU 10-423 L8–L9: Variational Inference, VAEs and the Diffusion ELBO

VAEs and diffusion models get stuck in the same place: log p_θ(x) requires integrating over latent variables, which is intractable. L8–L9 answer with variational inference. Pick a tractable q to approximate the true posterior, and swap 'minimize the KL' for 'maximize the ELBO,' which is a lower bound on log p(x). Add Monte Carlo estimation and the reparameterization trick, and a VAE trains with one forward and one backward pass. Unpack DDPM's ELBO and every term asks the learned reverse step to match the closed-form q(x_{t−1} | x_t, x₀).