Skip to content
Series
23 posts

Reading CMU 11-868 LLM Systems

A lecture-by-lecture reading of CMU 11-868 LLM Systems (Spring 2026) through its 28 public slide decks and seven MiniTorch assignments, from CUDA kernels and a homemade framework to distributed training, serving, and RLHF, with the no-video, GPU-required limits for self-learners noted throughout.

Reading CMU 11-868 LLM Systems: Overview and Self-Study Paths — 28 Slide Decks and 7 Assignments Are Public, but No Videos and You Bring Your Own GPU

CMU 11-868 is Lei Li's graduate course on LLM systems: it goes from CUDA kernels and your own MiniTorch framework to distributed training, SGLang serving, and RLHF. All 28 Spring 2026 slide decks, 7 assignment pages, and 7 starter-code repos are public, which rates it A3. What's missing: videos, GPUs and a PSC account, the quizzes, and any official statement of which two assignments are optional.

CMU 11-868 L01: Why LLMs Need Systems — the Scale Curve, Low-Level Operators, and Three Layers of Abstraction

CMU 11-868's first lecture spends 51 slides on one argument: the LLM bottleneck isn't only the model, it's computing larger LLMs on bigger datasets with fewer GPUs, less memory, and less power, faster. It breaks a Transformer into four low-level operators (matrix multiply, reduction, map, memory movement), sorts the hard problems into kernel, framework, and distributed-system layers, and warns that fast computation isn't enough because moving data takes time too.

CMU 11-868 L02–L04 GPU Programming and Acceleration: Threads, Blocks, the Memory Hierarchy, and Tiling

CMU 11-868's three GPU lectures answer one question: why does a correct CUDA matmul use only 2.48% of an A100's FP32 compute? L02 covers SMs, warps, and the grid/block/thread hierarchy. L03 covers cudaMalloc, cudaMemcpy, and kernel indexing. L04 uses tiling, coalesced access, and bank-conflict avoidance to bring data closer than global memory, which sits about 500 cycles away.

CMU 11-868 Assignment 1: Writing MiniTorch's map, zip, reduce, and matmul in CUDA

The first 11-868 assignment has you write four CUDA kernels in src/combine.cu (map 15, zip 25, reduce 25, matmul 30 points), wire them into MiniTorch's Python backend, and finish with a 5-point integration test. Shared-memory optimizations for reduce and matmul are marked Optional. The assignment page says plainly that you need a GPU, and grading uses private test cases.

CMU 11-868 L05: How a Deep Learning Framework Computes Gradients from a Computation Graph

L05 follows a small sentiment classification network throughout. It expresses computation as a graph, evaluates it in topological order, sends gradients back with the chain rule and vector-Jacobian products, and then takes apart TensorFlow v1's placeholder, variable, operation, and session. One slide is labeled "important for HW2".

CMU 11-868 Assignment 2: Implementing Autodiff in MiniTorch and Training a Sentiment Classifier

The second 11-868 assignment has three parts: autodiff's topological_sort and backpropagate (40 points), a Linear layer and MLP network (30), and binary cross entropy plus the training loop (30). You then train a sentiment classifier on SST-2 with GloVe embeddings and must reach 75% validation accuracy. The default backend is the CUDA kernels from Assignment 1, and the repo merged small Fall 2026 fixes on 2026-09-02.

CMU 11-868 L06–L07: Reading Transformers, T5, LLaMA, and GPT-3 Like a Systems Engineer

11-868 spends only two lectures on the model itself. L06 breaks the Transformer into embeddings, multi-head attention, FFN, LayerNorm, and residuals; L07 uses T5, LLaMA, and GPT-3 to show what modern LLMs changed. For a systems engineer the point is to remember the shapes: GPT-3 175B has 96 layers, d_model 12288, a 2048-token context, and trained on 300B tokens; LLaMA 65B has 80 layers, d_model 8192, and trained on 1.4T tokens. Those numbers set the workload for every acceleration, parallelism, and serving lecture that follows.

CMU 11-868 L08–L09: Choosing a Vocabulary, Emitting Tokens, and Why Speculative Decoding Is Fast

L08 goes from BPE to VOLT, a method co-authored by the lecturer Lei Li: vocabulary size has both a cost and a value, and VOLT finds the sweet spot by asking how much normalized entropy each added token removes, then solves it as an optimal transport problem. The second half covers LLaMA 3 growing its vocabulary from 32k to 128k and the cost of byte-level BPE splitting one Chinese character into three tokens. L09 moves from greedy decoding, sampling, and beam search to speculative decoding: a small model guesses N tokens and the big model checks them in one forward pass, because checking is cheaper than generating. It ends with EAGLE, which predicts final-layer features instead of tokens.

CMU 11-868 HW3: Build GPT-2 in Your Own MiniTorch and Make It Translate German

HW3 has you add softmax loss, Dropout, LayerNorm, and Embedding to the MiniTorch you built in HW1 and HW2, assemble a pre-LN GPT-2 decoder, and train it on IWSLT14 German-English translation. Points: tensor functions 20, basic modules 20, decoder LM 40, translation pipeline 20. Full marks require passing the private tests and a BLEU of about 20±2. The assignment page warns that training alone takes at least 10 hours: one epoch is about an hour on a PSC V100, and you need 10. In Spring 2026 it went out Feb 4 and was due Feb 18.

CMU 11-868 L10: Accelerating Transformers on GPUs, and Where LightSeq Finds the Time

Lecture 10 of 11-868 uses Lei Li's own LightSeq and LightSeq2 as the case study and breaks them into four techniques: fuse every small operation outside matrix multiplication into one kernel, rewrite the LayerNorm and Softmax formulas to cut thread synchronizations, store parameters and gradients in FP16 but compute updates in FP32, and reuse memory based on backward-pass dependencies. The slides report 1.4-3.5x training speedups on WMT14 English-German. There is no recording; this guide works from slide page numbers and the two papers.

CMU 11-868 HW4: Writing Softmax and LayerNorm as Fused CUDA Kernels

The fourth 11-868 assignment has you follow LightSeq and hand-write CUDA kernels for attention softmax and LayerNorm (forward and backward), bind them into your own MiniTorch, then swap them into your HW3 Transformer and train for one epoch. Points: Softmax 40, LayerNorm 40, integration 20. The assignment page expects individual kernels to be 3.7x to 15.8x faster, but end-to-end training only about 1.1x faster, because of Amdahl's law. You need an NVIDIA GPU, and the repo has already been changed for Fall 2026.

CMU 11-868 L14-L15: Distributed Training and Data Parallelism, and Where Gradient Sync Costs Come From

Lectures 14 and 15 of 11-868 go from the parameter server to PyTorch DDP. They use NCCL's five collectives (Broadcast, Reduce, AllReduce, ReduceScatter, AllGather) as building blocks, show why a ring makes broadcast time nearly independent of GPU count, and split AllReduce into ReduceScatter plus AllGather. The second lecture takes apart DDP's two key designs: bucketing gradients (25 MB by default) and starting synchronization before the backward pass finishes. There is no recording; this guide works from slide page numbers and the VLDB 2020 paper.

CMU 11-868 L16–L17: When a Model Won't Fit on One GPU — Split Layers, Matrices, or Experts

CMU 11-868 (Spring 2026) spends two lectures on models too big for one GPU. L16 covers pipeline parallelism, which splits layers (GPipe micro-batches, 1F1B, interleaved stages), and tensor parallelism, which splits matrices (Megatron-LM's cuts for FFN, attention, and embeddings). The rule of thumb: TP inside a node, PP across nodes, DP on top. L17 treats MoE as a third way to split: each GPU holds different experts and replicates everything else. The price is all-to-all communication and load balancing, shown through GShard, DeepSpeed-MoE, and DeepSeek-V3.

CMU 11-868 L18: How ZeRO Cuts Data-Parallel Memory — Optimizer State, Gradients, Then Parameters

Data parallelism keeps a full copy of parameters, gradients, and optimizer state on every GPU. With Adam and mixed precision that is about 20 bytes per parameter, 16 of them optimizer-related, so LLaMA-3 8B already needs 160GB. CMU 11-868 L18 builds on the ZeRO paper and animates its three stages frame by frame: ZeRO-1 partitions optimizer state, ZeRO-2 also partitions gradients, and ZeRO-3 partitions parameters too. The slides conclude that the first two stages add no communication and save up to 8x memory; stage 3 makes per-GPU memory shrink with GPU count, at what the slides estimate as about 3x the communication.

CMU 11-868 HW5: Writing Data Parallelism and Pipeline Parallelism Yourself on Two GPUs

CMU 11-868's fifth assignment switches to PyTorch and Hugging Face GPT-2. Using only torch.distributed and torch.multiprocessing, you write data parallelism (partition the data, set up a process group, average gradients; 50 points), then a GPipe-style pipeline (split the model, generate a clock schedule, run micro-batches on worker threads; 50 points). Both parts need benchmarks and plots on at least two GPUs: data parallelism must reach at least 1.5x speedup on 2 GPUs, and the pipeline must beat plain model parallelism. The Spring 2026 deadline was 3/25.

CMU 11-868 L19–L20 Model Quantization: GPTQ Saves Memory, and the Speedup Comes Along for the Ride

11-868 spends two lectures on quantization. L19 goes from BF16 and absmax/zero-point quantization to AdaQuant, ZeroQuant, and LLM.int8(). L20 is all GPTQ. GPTQ quantizes weights only: after each column is quantized, it uses second-order information to adjust the weights not yet quantized, and lazy batch updates plus a Cholesky trick let it scale to 175B. What it mainly saves is memory. Inference gets faster because single-batch decoding was already bottlenecked on reading weights; the amount of arithmetic does not shrink.

CMU 11-868 L21 FlashAttention: Attention Is Slow Because of Data Movement, Not Math — Tri Dao from FA1 to FA4

Standard attention writes the N×N score matrix out to HBM and reads it back, and most of its time goes to that traffic. FlashAttention uses tiling plus softmax rescaling so each block finishes inside SRAM, and the backward pass recomputes instead of storing. Tri Dao's guest slides for 11-868 give one set of numbers: the backward pass does 13% more FLOPs, 9x less HBM traffic, and runs 6x faster. FA3 and FA4 follow the same theme: when the hardware changes, the bottleneck moves, and the algorithm has to move with it.

CMU 11-868 L12–L13 TPU, JAX, and Pallas: One Attention Kernel on TPU, from XLA Fusions to Splash Attention

Across two lectures and more than 200 slides, Google's Srinath Mandalapu traces one attention computation from Python down to TPU VLIW instructions. L12 covers the JAX ecosystem, the memory and compute units of TPU Ironwood, and how XLA compiles attention into three fused kernels. L13 covers what XLA cannot do: using Pallas to control movement between HBM and VMEM yourself, writing FlashAttention, then adding block sparsity to get Splash Attention. The ideas match the GPU version. The difference is that on TPU the compiler does most of the scheduling, and Pallas is how you take loops and block sizes back into your own hands.

CMU 11-868 L23 Efficient Fine-Tuning for Large Models: LoRA, CIAT, and QLoRA

Lecture 23 of 11-868 treats fine-tuning as a memory problem. Full-parameter half-precision fine-tuning of LLaMA-8B needs about 80GB. LoRA brings that to about 33GB, and QLoRA, which stores the frozen weights in 4 bits, gets it to about 9.2GB. The lecture moves in three steps: train only two small low-rank matrices A and B (the slides credit CIAT as the first to do this); squeeze frozen weights to about 0.52 bytes per parameter with an NF4 lookup table and double quantization; and use a paged optimizer to push optimizer state to the CPU when the GPU is about to run out.

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

CMU 11-868 HW6: Training Llama-2-7B with DeepSpeed ZeRO + LoRA, Serving with SGLang

The sixth 11-868 assignment is the first to set aside your homemade MiniTorch and use industry frameworks. Two problems, 50 points each. Problem 1: edit a DeepSpeed training script to turn on LoRA so Llama-2-7B can train on two 16GB V100s. Problem 2: fill in the TODOs of an SGLang inference script and tune parameters to make generation faster. The two problems want conflicting GPUs: SGLang doesn't support V100, so you need an L40S, A6000, or A100. The spring due date was April 13, and the assignment page publishes no grading tests.

CMU 11-868 Serving at Scale: Prefill/Decode Disaggregation, KV Cache, and Heterogeneous Hardware

11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.

CMU 11-868 RLHF Systems and Assignment 7: A VERL-Style Pipeline with a Reward Model, GAE, and PPO

11-868's RL systems lecture has no slides; the Syllabus lists just one paper, ReaLHF. Assignment 7, on the other hand, is fully public. You train a DistilBERT reward model on Anthropic's HH-RLHF data (40 points), fill in GAE, the PPO loss, and entropy in a VERL-style trainer to fine-tune GPT-2 (40 points), and compare reward distributions before and after RLHF (20 points). The starter trainer never imports the verl package. What you learn is the RLHF dataflow, not VERL's distributed engine.