🌏 中文版
Edition note: This series follows the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is the latest complete term in 2025–2026: the last schedule entry is the April 30 final report deadline, and the footer reads "Last updated April 20, 2026."
10423-f26/returns 404, so there is no Fall 2026 site. Every fact was checked on 2026-09-30 against the course homepage, the schedule, the Coursework and Previous pages, and the slide and homework files. Access rating: A3 (defined in the global AI/CS course map).
Series position: this is the overview | Next: L1: RNN language models and autodiff (with HW0)
CMU 10-423 is the generative AI course in Carnegie Mellon's Machine Learning Department. In Spring 2026 it was co-taught by Aran Nayebi and Matt Gormley. One class carries three course numbers, 10-423, 10-623 and 10-723. The content is identical; what you have to turn in differs.
The course description is broad. It covers how to build generative models and other large foundation models (transformers for vision and language, diffusion models), how to train them (pre-training, fine-tuning) and adapt them efficiently (adapters, in-context learning), how to scale to massive datasets (multi-GPU and distributed optimization), how to use existing models day to day (generating code, coding with a generative model in the loop), and what can go wrong (bias, hallucination, adversarial attacks, data contamination) along with ways to fight it.
This post answers three questions: what the course teaches, what an outside reader can actually get, and how to study it on your own. Lecture details come in the later posts.
The hard facts
| Item | Spring 2026 |
|---|---|
| Instructors | Aran Nayebi, Matt Gormley |
| Meetings | MWF 2:00–3:20 PM (DH 2210); lectures on Mondays and Wednesdays, occasional recitations on Fridays |
| Prerequisites | One of 10301, 10315, 10601, 10701, 10715, 11485, 11685, 11785 |
| Textbook | None; readings are free papers and book chapters online |
| Homework language | Python |
| Past sites | Fall 2025, Spring 2025, Fall 2024, Spring 2024 (Previous page) |
The prerequisite list has two tracks: intro machine learning (10-301/315/601/701/715) or intro deep learning (11-485/685/785). Lecture 1 adds a note: deep learning and PyTorch are not required. Depending on which prerequisite you took and when, you may or may not have seen them, and either way is fine.
This site has guides to both tracks: Reading CMU 10-301 and Reading CMU 11-785.
How the three course numbers differ
The syllabus is blunt: the three courses are identical in content, except that 10-623 students also do HW623 and 10-723 students do HW623 and Quiz723.
| Course | Homework | Quizzes | Programming tests | Exam | Project | Participation |
|---|---|---|---|---|---|---|
| 10-423 | 30% (5 total) | 10% (6, lowest at half weight) | 10% (2) | 20% (1) | 25% | 5% |
| 10-623 | 30% (6 total, incl. HW623) | 10% (same) | 10% | 20% | 25% | 5% |
| 10-723 | 30% (6 total) | 10% (6, equal weight) | 10% | 20% | 25% | 5% |
"5 homeworks" means HW0 through HW4. According to the homework table in Lecture 1, HW623 asks you to read and analyze a recent generative AI paper and present it on video.
Two pieces of official material disagree with each other. Go by the homepage syllabus:
- Lecture 1's "Syllabus Highlights" slide says "40% homework" with no programming tests, and its "Reminders" slide lists HW0 as out on August 27. Those look carried over from an earlier term. Lecture 2 already has the Spring 2026 dates (out January 14, due January 26), matching the schedule.
- The Coursework page says "There will be 5 quizzes," then lists Quiz 1 through Quiz 6. The syllabus says 6.
Homework policy: yourself first, AI second
The course's most distinctive design is the two-stage homework submission. The syllabus's reasoning: the most important learning in the course happens while you struggle through hard homework problems.
- Slot A (human work only): individual work with limited collaboration and no AI assistance of any kind. Staff grade it and tell you which questions you missed. Office hours and Piazza are only active during Slot A.
- Slot B (AI assistance and full collaboration allowed): due three days after you get feedback. Only the questions you missed in Slot A are regraded.
- Score: each question keeps the higher of the two scores, with a bonus for scoring above half in Slot A.
The full formula from the syllabus
$$ s = 0.95 \times \sum_{q \in HW} \max(s_{A,q}, s_{B,q}) + 0.05 \times \mathbb{1}(s_A > 0.50) $$
$s_{A,q}$ and $s_{B,q}$ are your scores on question $q$ in each slot, and $s_A$ is your Slot A total.
Submitting AI-assisted work to Slot A as human work counts as an academic integrity violation, with penalties up to failing the course. Late rules apply to Slot B only. Slot A accepts no late work and no grace days. Slot B loses 25% per day late, down to 25% credit on day three, and you get 6 grace days for the semester.
Lecture 1 spends several slides defending this design, including one candid concession. Yes, these problems are easy for an LLM or a coding agent. But to use a code assistant well on a problem nobody has solved, you need to read lots of generated code, spot bugs in code that looks correct, and state clearly what the problem is.
What this means for self-learners: nobody will grade your Slot A, but you can keep the same order. Do a full pass without AI, find your mistakes against the practice exam solutions or the unit tests, then bring in AI to fix them.
Map of the 6 units and 26 lectures
The schedule groups the 26 lectures into 6 units, and each homework covers only the lectures before it: HW1 covers L1–L4, HW2 L5–L8, HW3 L9–L12, HW4 L12–L14. This series follows that rhythm and closes each unit with a homework post.
| Unit | Lectures | Posts in this series |
|---|---|---|
| Generative models of text | L1 RNN LMs / Autodiff; L2 Transformer LMs; L3 Learning LLMs / Decoding; L4 Pre-training, fine-tuning / Modern Transformers | 1 RNN LMs and autodiff, 2 Transformer LMs and decoding, 3 Modern Transformers, 4 HW1 |
| Generative models of images | L5 CNNs / Encoder-only Transformers / ViT; L6 GANs / PGM; L7–L8 Diffusion models | 5 CNN/BERT/ViT, 6 GANs, 7 Diffusion, 8 VI and VAEs, 9 HW2 |
| Applying and adapting foundation models | L9 VAEs; L10 Parameter-efficient fine tuning; L11 In-Context Learning / Prompt Engineering / Instruction Fine-tuning / RLHF | 10 PEFT and ICL, 11 IFT/RLHF/DPO, 12 HW3 |
| Multimodal foundation models | L12 DPO / Text-to-image / Latent diffusion; L13 Vision-language models; L14 Cross-Attention / DiT / Prompt-to-Prompt | 13 Text-to-image and VLMs, 14 Cross-attention/DiT/Q-Former, 15 HW4 |
| Scaling Up | L15 Querying Transformer / Scaling Laws; L16 Mixture of Experts; L17 Distributed training; L18 Flash Attention / Efficient decoding | 16 Scaling laws and MoE, 17 Distributed training and efficient inference |
| Advanced Topics | L19 Long Context; L20 Reasoning Models; L21 State Space / Hybrid Models; L22 Real-world Issues; L23 Code Generation / Autonomous Agents; L24 Audio; L25 Video; L26 Interactive World Models + Science of Alignment | 18 Long context and SSMs, 19 Reasoning models, 20 Risks and alignment, 21 Code generation and agents, 22 Audio, video and world models |
The last post, 23 Practice exam, HW623 and the final project, wraps up.
Three lecture decks span two topics: L9 is named vae-icl, L12 dpo-text2img, and L15 querying-scaling. This series splits each by topic across two posts, so posts and lectures do not map one to one.
Assessment timeline
| Date (2026) | Event |
|---|---|
| 1/14 | HW0 out |
| 1/26 | HW0 Slot A due, HW1 out |
| 1/28 | Quiz 1 (L1–L4) |
| 2/9 | HW1 Slot A due, HW2 out |
| 2/16 | Quiz 2 (L5–L9) |
| 2/21 | HW2 Slot A due, HW3 out |
| 2/25 | Programming test HW1/HW2, Quiz 3 (L9–L12) |
| 3/12 | HW3 Slot A due, HW4 out |
| 3/16 | Quiz 4 (L12–L15) |
| 3/23 | HW4 Slot A due, HW623 and practice problems out |
| 3/27 | Programming test HW3/HW4 |
| 3/30 | Exam (evening) |
| 4/3, 4/13 | Project proposal and midway report due |
| 4/6, 4/20 | Quiz 5 (L16–L20) and Quiz 6 (L21–L24); HW623 due 4/20 |
| 4/26–4/30 | Project poster due, final presentations, final report due |
There are no programming homeworks after L15. The second half is assessed through Quizzes 5–6, HW623 and the project, which is done in teams of three over the last four weeks.
Access rating: A3, with the gaps spelled out
The course rates A3 because the core self-study material is all there. You get slides for all 26 lectures (the schedule links 40 PDFs; 13 are inked versions from class, and L26 has two decks), HW1–HW4 zips (handout PDF, starter code, unit tests, LaTeX template), a read-only Overleaf template per homework, the HW623 handout, a practice exam with solutions, and the project handout.
The practice exam is the Spring 2026 version: 41 pages, 167 points, 13 questions, running from AutoDiff/RNN-LMs to Scaling Laws. It works well as a self-check for each unit.
What was unavailable when tested on 2026-09-30:
| Material | Status | Effect |
|---|---|---|
| Lecture and recitation recordings | On SCS Panopto; the schedule says "Andrew ID Required" and asks you to sign in through Canvas; the livestream link is on Piazza | The whole series is written from slides |
| Past recordings | The four past sites (F25, S25, F24, S24) also link only to Panopto, with no public videos | No older videos to fall back on |
| HW0 PyTorch Primer handout | Google Drive link returns 401 | Only the public HW0 recitation Colab is readable |
| HW1 recitation slides | Public Google Slides | Readable |
| HW1 Supplemental Material | Drive returns 401 | Missing |
| HW2 recitation slides | Public Google Slides | Readable |
| HW3, HW4 recitation slides | Return 401 | Flagged in the homework posts |
| Whiteboard notes | The schedule mentions a OneNote notebook, but the link is empty | Missing |
| Gradescope, 6 quizzes, 2 programming tests, the real exam, Piazza | Enrolled students only | No grading; self-assess with the practice exam solutions and unit tests |
Also, from L9 onward most schedule entries list no readings. This series leaves that as is and cites only the slides, without inventing reading lists.
Compute needs for the four homeworks
| Homework | Topic (per Lecture 1) | Verified environment notes |
|---|---|---|
| HW0 | PyTorch Primer: image and text classifiers | Handout returns 401; the recitation Colab covers PyTorch, LSTMs, Weights & Biases and einops |
| HW1 | Large Language Models: add GQA and RoPE to a Transformer LM | Handout due 2/9, 62 points; includes setup notes for Colab (free T4) and Kaggle (30 free hours of T4/V100 per week) |
| HW2 | Image Generation: diffusion model | Checked separately in the HW2 post |
| HW3 | Adapters for LLMs: GPT-2 + LoRA | Handout due 3/12, 66 points; ships run_in_colab.ipynb and wandb_api.json; the handout notes Colab's free T4 |
| HW4 | Multimodal Foundation Models: text-to-image | Handout due 3/23, 79 points; ships download_data.sh and run_in_cloud.ipynb; the handout recommends claiming Colab Pro with a CMU email and using an A100 when available |
The Colab Pro tip for HW4 needs a CMU email, so outside readers will have to find their own GPU. Details are in each homework post.
A self-study path
If you only want the concepts: read the slides in schedule order and, after each unit, attempt the matching practice exam questions. You can follow without doing the homework, but you will miss the heaviest part of the course.
If you want to do the homework:
- Run the HW0 recitation Colab first to make sure PyTorch and W&B work for you.
- After each unit's slides, do the matching homework: HW1 → HW2 → HW3 → HW4. Write the written parts yourself first, Slot A style, and check the programming parts with the unit tests in the zip.
- After all four homeworks, take the practice exam, then check the solutions.
- The second half (L15–L26) has no homework. Pick a topic and run a small project in the format of the project handout.
One thing to do tonight: open the schedule, download the L1 slides, and run the first cell of the HW0 recitation Colab.
Further reading
This course overlaps with several series on this site. Every post in this series stands on its own; the links below are for going deeper:
- Building language models from scratch, scaling, parallelism and inference: Reading Stanford CS336
- Transformers, LLM training, preference tuning and agents: Reading Stanford CME295
- The math of diffusion and flow matching: Reading MIT 6.S184
- The systems side of LLMs (CUDA, distributed training, serving): Reading CMU 11-868
- Prerequisites: Reading CMU 10-301, Reading CMU 11-785
References
- CMU 10-423/623/723 Generative AI homepage (Spring 2026): course description, learning outcomes, prerequisites, grading, Slot A/B policy, late rules
- Schedule: 26 lectures, 6 units, readings, homework and quiz dates, Panopto notes
- Coursework: HW0–HW4, HW623, quiz coverage, practice exam, project milestones
- Previous Course Homepages: four past sites from Fall 2025 back to Spring 2024
- Lecture 1 slides: Course Overview + RNN-LMs + Automatic Differentiation: homework table, prerequisite notes, the case for the homework policy
- Lecture 2 slides: Transformer Language Models: Spring 2026 HW0 dates
- HW1 handout (zip), HW3 handout (zip), HW4 handout (zip): due dates, point totals, compute notes
- Practice Exam and Solutions: Spring 2026, 41 pages, 167 points, 13 questions
- Project handout
- HW0 recitation Colab
Loading...