Skip to content

Reading CMU 11-868 LLM Systems: Overview and Self-Study Paths — 28 Slide Decks and 7 Assignments Are Public, but No Videos and You Bring Your Own GPU

Sep 30, 20261 min
TL;DRCMU 11-868 is Lei Li's graduate course on LLM systems: it goes from CUDA kernels and your own MiniTorch framework to distributed training, SGLang serving, and RLHF. All 28 Spring 2026 slide decks, 7 assignment pages, and 7 starter-code repos are public, which rates it A3. What's missing: videos, GPUs and a PSC account, the quizzes, and any official statement of which two assignments are optional.

🌏 中文版

Version note: This series follows the Spring 2026 offering of CMU 11-868 LLM Systems, the most recent completed term with the fullest materials. Fall 2026 is in progress and is used only for comparison. Every fact was checked on 2026-09-30 against the official pages, slide PDFs, and GitHub repos. Access rating: A3 (defined in the global AI/CS course map). Slides, assignment specs, starter code, and the project spec are public, which is enough to self-study. Videos, the GPU cluster, quizzes, and grading are not.

Series: this is the overview | next L01: Why LLMs Need Systems

CMU 11-868 is a graduate course in the Language Technologies Institute, taught by Lei Li. It doesn't teach you how to use LLMs. It teaches you how to turn LLM training, fine-tuning, and serving into systems that actually run. The Spring 2026 course description lists eight topics: training efficiently on huge data, embedding storage and retrieval, data-efficient fine-tuning, communication-efficient algorithms, efficient RLHF, acceleration on GPUs and other hardware, model compression for deployment, and online maintenance.

The homework centers on MiniTorch. The assignment site overview says it began as a teaching framework by Sasha Rush, and the course will "extend the original framework to support real cuda kernels." You write CUDA kernels first. Then you build autodiff, a Transformer, and fused kernels inside your own framework. Only near the end do you switch to industry frameworks like DeepSpeed and SGLang.

The FAQ draws a clear line between this course and CMU's other LLM course, 11-667. 11-667 covers models, learning algorithms, and applications. 11-868 covers "building systems for LLM, including training, serving, and maintaining." The FAQ also says plainly that students who don't want to write low-level systems code should take 11-667.

The hard facts

ItemSpring 2026
Schedule1/12–4/28, Mon/Wed 12:30–1:50, optional Friday recitations
StaffLei Li; four TAs (Aditya Tummala, Danqing Wang, Jackey Hua, Sreeram Vennam)
PrerequisitesLinear algebra, calculus, probability and statistics; Python plus C/C++/Java (15-122 level); ML background "preferred but not required"
Homework"five required and two optional programming assignments", done individually
GradingHomework 44% (+5% optional), Quiz 10%, Participation 2% (+ up to 5%), Project 44%
TextbookNone required; Programming Massively Parallel Processors, 4th ed., recommended for anyone new to GPU programming
ComputePSC cluster; Google Colab for new users
ForumEd; don't email individual TAs
VideosNo video or YouTube link anywhere on the official pages

Sources: Logistics, Syllabus, FAQ.

The FAQ estimates about 12 hours a week for students who have taken an ML course and are comfortable in C. The participation bonus fits the course well: find a typo or bug in an assignment repo, send a pull request, and get credit if it's merged. The latest commits on llmsys_hw1 through llmsys_hw7 are mostly "Merge pull request," so people do use this channel.

How four terms evolved

The course hub lists four terms: Spring 2024, Spring 2025, Spring 2026, and Fall 2026. It also notes that a "slightly adjusted" version is a core course in the CMU GenAI/LLM certificate.

TermHomeworkAssessmentSchedule changes
Spring 2024HW1–HW410% per assignment; in-class paper presentation 10%; every student writes paper reviewsLate-term lectures on RAG, HNSW, multimodal LLMs, attention sinks
Spring 2025Syllabus lists HW1–HW5 due datesLogistics says "Homework 10% each, 40% in total"; $150 of AWS credit per studentGuest lectures by Tri Dao, Woosuk Kwon, Hao Zhang, Ying Sheng, and others; last lecture on RL systems
Spring 2026HW1–HW7 (5 required + 2 optional)Homework 44% (+5%)Google guest lectures on TPU/JAX and Pallas; guests on disaggregated prefill/decode (Vikram Mailthody) and LMCache (Junchen Jiang)
Fall 2026Same 7 assignmentsHomework 44%, "+5% optional" droppedNew Week 13 "Acceleration on TPU 1/2"

The two Spring 2025 pages disagree: the Syllabus schedules five assignments, while Logistics says 10% each and 40% in total. The course doesn't explain this, and this article doesn't guess.

The overall direction is clear. Spring 2024 still ran partly as a seminar, with paper reviews and student presentations. Those were dropped in favor of more programming assignments and industry guests. The late-term lectures on RAG, vector search, and multimodal models moved to an "unscheduled" block at the bottom of the Spring 2026 Syllabus, with readings but no slides.

A3, with six gaps

Spring 2026 publishes a lot: 28 slide PDFs on the Syllabus, seven specs on the assignment site, seven assignment repos plus the example code in llmsys_code_examples under the GitHub org, and the project spec. By the course map's definition, that's A3: structured materials plus assignments and the files they need.

A3 still has gaps. Outside readers will run into these six:

  1. No videos. Neither the Spring nor the Fall 2026 pages link to recordings. This is the biggest difference from Stanford CS336. You can only read slides and papers, and anything said out loud in class is lost.
  2. You need NVIDIA GPUs. Assignment 1 opens with "You'll need a GPU." Assignment 5 needs at least two. Assignment 6 recommends a 2-GPU PSC session and asks you to make Llama-2-7B trainable with LoRA on "2 V100 GPUs/ 16GB GPU memory." Without a PSC account, you rent cloud GPUs yourself.
  3. Quizzes and grading are closed. Quizzes are 10% of the grade, and the quiz links on the slides point to CMU Canvas. Ed and the submission systems are also for enrolled students only.
  4. The course never says which two assignments are optional. Logistics only says "five required and two optional," and none of the seven assignment pages is marked optional. This series won't guess.
  5. Some lectures have no slides. The 4/15 lecture "Efficient Reinforcement Learning System for LLMs" only links the ReaLHF paper. Parts 1 and 2 of "Accelerating Transformer on GPU" share one PDF. Five unscheduled topics at the bottom of the Syllabus (Triton, RAG, HNSW, multimodal LLMs, attention sinks) have readings only.
  6. You buy the textbook. The PMPP 4th edition link goes through O'Reilly's CMU SSO, and Logistics says "Please login using your andrew email for free access."

The FAQ also says there's no audit option, because the waitlist comes first.

Assignment repos are shared across terms

This affects the code you clone. The assignment site and llmsys_hw1–llmsys_hw7 aren't split by term, and Fall 2026 is editing them now. Latest commit on the default main branch, via the GitHub API on 2026-09-30:

RepoLatest commitStatus
llmsys_hw12026-01-30During spring
llmsys_hw22026-09-02Changed for Fall 2026
llmsys_hw32026-09-09Changed for Fall 2026
llmsys_hw42026-09-30Changed for Fall 2026
llmsys_hw52026-04-30Spring state
llmsys_hw62026-03-23Spring state
llmsys_hw72026-05-02Spring state

The older llmsys_s24_hw1–4 and llmsys_s25_hw1–5 repos are still public.

What to do: to match the spring version, check out the last commit before the term ended:

git clone https://github.com/llmsystem/llmsys_hw2.git
cd llmsys_hw2
git checkout $(git rev-list -n 1 --before=2026-04-29 main)

To pick up Fall 2026 fixes, stay on main, but problem numbers and points may differ from what this series describes.

How the other 22 posts are ordered

The series mostly follows the official schedule with three changes, all made for readability rather than by the course: each assignment post comes right after the lectures it depends on (the course often releases homework first; HW1 comes out on the day of L02); Google's TPU/JAX and Pallas guest lectures move from Week 7 to after FlashAttention, so Splash Attention can be read against FlashAttention's tiling; and the three unscheduled serving decks with slides are folded into the serving-at-scale post.

#StagePost
1MotivationL01 Why LLMs need systems
2HardwareL02–L04 GPU programming and acceleration
3HardwareHW1: CUDA Programming
4FrameworkL05 Deep learning frameworks and autodiff
5FrameworkHW2: MiniTorch Framework
6ModelL06–L07 Transformers and pre-trained LLMs
7ModelL08–L09 Tokenization, decoding, and speculative decoding
8ModelHW3: A decoder-only Transformer in MiniTorch
9Single-GPU speedL10 Accelerating Transformers on GPU (LightSeq)
10Single-GPU speedHW4: Fused Softmax/LayerNorm kernels
11Multi-GPU trainingL14–L15 Distributed training and data parallelism
12Multi-GPU trainingL16–L17 Model parallelism and MoE
13Multi-GPU trainingL18 ZeRO
14Multi-GPU trainingHW5: Data and pipeline parallelism
15Smaller and fasterL19–L20 Model quantization
16Smaller and fasterL21 FlashAttention
17Smaller and fasterL12–L13 TPU, JAX, and Pallas
18Smaller and fasterL23 Efficient fine-tuning (LoRA, QLoRA)
19ServingL22, L24 LLM serving: SGLang and vLLM
20ServingHW6: DeepSpeed + SGLang
21ServingL26–L30 Serving at scale and KV caches
22Alignment systemsRLHF systems and HW7

What Fall 2026 changed

ItemSpring 2026Fall 2026
ScheduleMon/Wed 12:30–1:50Mon/Wed 5:00–6:20, Friday recitation; Silicon Valley students join online
TAs46
AI tools"Using Github copilot or any AI agent is ok for explanation purpose""It is strictly forbidden to use any AI agent to complete the homework"; LLMs for explanation only
Late work3 free late days per term, then 20% off per day; no late final report"Each late day will incur 30% discount on the grades"
PSCNo guarantee when a job startsNormal waits exceed 24 hours; no extension because a job didn't run (unless PSC is down for more than 24 hours)
Homework weight44% (+5% optional)44%
Schedule orderGoogle guests in Week 7Google guests moved to Week 6; new Week 13 "Acceleration on TPU 1/2", no slides yet

Sources: Fall 2026 Logistics, Fall 2026 Syllabus.

The Fall 2026 course description was also rewritten. It names vLLM, SGLang, and FlashAttention, and says students will write "custom GPU/TPU acceleration kernels (CUDA/Triton)." On 2026-09-29 the GitHub org created a shared-tpu-notebooks repo, described as giving hundreds of students a Cloud TPU from a Jupyter notebook. The official pages don't say whether it's for the Week 13 TPU lectures.

Final project: spec and timeline

The project is 44% of the grade: proposal 2%, mid-term 2%, presentation 20%, final report 20%. From the Projects page:

  • Teams: 2–3 people. The topic needs both a systems side and an LLM side.
  • Two project types: reimplement a recent LLM systems paper inside MiniTorch, or do research aimed at a venue like MLSys, OSDI, SOSP, or SC. The L01 slides add that MiniTorch projects "should not use external PyTorch/Tensorflow code."
  • Proposal: MLSys 2024 LaTeX style. Cover the systems problem, the state of the art, evaluation and workload, team split, timeline, and how much CPU/GPU/storage and compute time you need.
  • Mid-term report: no page limit, suggested max 6 pages: motivation, related work, method, early experiments, remaining work.
  • Final report: suggested max 8 pages. It adds implementation details, analysis and ablations, limitations, and file-level member contributions (e.g., "Author X contributed to the function in the file ABC.py").
  • Seed topics: FlashAttention, PagedAttention, mixed-precision training, or DPO in MiniTorch. Research ideas include KV cache management, faster MoE training, training on heterogeneous hardware, and making WebLLM faster in the browser.

Spring 2026 timeline (Syllabus):

MilestoneDate
Team list2/20
Proposal2/27
Mid-term report4/1
Final presentation4/27
Final report4/28

The L03 slides give 2/18 as the team deadline. The Syllabus says 2/20; go with the Syllabus.

Self-study paths

The materials are public. The real barrier is hardware. Check what you have, then pick a path.

Path 1: slides only, no GPU

Read the slides and the Syllabus readings in this series' order. The only assignment you can do is Assignment 2, which builds autodiff and a sentiment classifier in MiniTorch. Its page says "Training on CPU can take some time," so a CPU works.

Tonight: open L02 GPU Programming, copy the B200/H100/A100 spec table on page 15, and note the ratio of FP32 compute to memory bandwidth. Several later lectures use it.

Path 2: one NVIDIA GPU

Add HW1 (CUDA kernels), HW3 (a decoder-only Transformer in MiniTorch), and HW4 (fused Softmax and LayerNorm kernels). A Colab T4 is fine for the examples: the llmsys_code_examples notebook compiles with -arch=sm_75 for the T4. HW3's Problem 4 trains a translation model, and the page warns "You may need to spend at least 10 hours for the training process." Free Colab sessions won't last that long, so rent an hourly cloud GPU.

Tonight: start a Colab GPU runtime, run simple_cuda_demo/CUDA_Code_Examples.ipynb, and confirm nvcc compiles the vector-add example.

Path 3: two or more GPUs

HW5 (your own data parallel and pipeline parallel) needs at least two GPUs. HW6 (DeepSpeed ZeRO plus LoRA training on Llama-2-7B, then SGLang inference) targets two V100s. HW7 builds RLHF in a VERL-like framework; its commands use GPT-2, and the page doesn't state hardware requirements.

Tonight: price it out. Multiply a two-GPU machine's hourly rate by HW3-style "at least 10 hours" training runs, then decide whether to do the whole path or only the data-parallel half of HW5.

The project

You won't have teammates or a grader, but the project spec makes a good practice list. Pick a seed topic, such as FlashAttention in MiniTorch, and write it up for yourself in the five-section mid-term format.

Further reading

References