🌏 中文版
This post covers Lecture 4, "LLM training," of the 2025 edition of Stanford's CME295 (October 17, 2025). The main source is the 128-page slide deck; the recording is there to watch alongside. Everything below is based only on the text and figures on the slides, not on anything said out loud in class.
The first three lectures answered "what does an LLM look like?" Lecture 4 asks a different question: how does that machine go from random weights to an assistant that answers questions? The slides give a two-part answer. First pretraining, which teaches the model the patterns of language and code. Then finetuning, which teaches it to do what it's told. The two stages differ in cost by several orders of magnitude, and most of the lecture is about where the money and memory go and how to save them.
Start with a counterexample: a pretrained model only continues text
The slides open with a question: "Can I put my teddy bear in the washer?"
A model that has only been pretrained replies: "Teddy bears are often made of materials like polyester and cotton, with plastic eyes and sometimes small accessories." Nothing in it is wrong. It is simply continuing with the kind of text a web page might contain, and never answers the question.
A model that has been instruction tuned replies: "No, it might get damaged. Try hand washing instead."
The gap between those two answers is the two stages of this lecture.
flowchart LR
I["Randomly initialized model"] -->|"Pretraining<br/>trillions of tokens, next-token prediction<br/>cost: millions of dollars at least"| P["Pretrained model<br/>knows language and code, but only continues text"]
P -->|"SFT / instruction tuning<br/>thousands to millions of demonstrations"| S["Assistant model<br/>follows instructions"]
P -.->|"LoRA / QLoRA<br/>train only a small fraction of parameters"| S
S -->|"Preference tuning (Lecture 5)"| A["Model that misbehaves less"]
The two-stage recipe is itself a paradigm shift. The slides contrast three approaches. Traditional machine learning trains a model from scratch for every task: one for spam detection, one for sentiment extraction, one for translation. Transfer learning reuses part of a trained model. LLM training first trains a model to understand language, then tunes it for the end task.
Stage one: pretraining
Objective and data
The pretraining objective is simple: predict the next token. Given "[BOS] A teddy bear is", guess the next word.
The goal is to learn "patterns of language and code," so the data mixes web-scraped text (the slides name Common Crawl and Wikipedia) with code (GitHub, Stack Overflow), across many languages. The scale is trillions of tokens:
| Model | Pretraining tokens |
|---|---|
| GPT-3 (2020) | 300 billion |
| Llama 3 (2024) | 15 trillion |
Measuring scale in FLOPs
The slides separate two terms that often get mixed up. FLOPs is the total amount of computation (how many floating-point operations were done); FLOPS or FLOP/s is how many can be done per second. The first is the training bill, the second is hardware speed. In orders of magnitude, training a small neural network takes about 10⁷ FLOPs, a large RNN about 10¹⁴, an LLM about 10²⁵. On the hardware side, a smartphone does about 10¹² FLOPS, a computer about 10¹⁴, a supercomputer about 10¹⁸.
Scaling laws: bigger models are more sample-efficient
The slides take two findings from Kaplan et al.'s scaling-law paper:
- Scaling: test loss falls as a power law in compute, dataset size, and parameter count, each a straight line on log-log axes
- Sample efficiency: larger models reach lower loss with the same number of tokens, so bigger models learn faster
Then comes the Chinchilla law. The slides reproduce the paper's table of "compute-optimal" training tokens for different model sizes:
| Parameters | Optimal training tokens |
|---|---|
| 1 billion | 20.2 billion |
| 67 billion | 1.5 trillion |
| 175 billion | 3.7 trillion |
Reading off the table, that's roughly 20 tokens per parameter. Put it next to the earlier table: GPT-3 has 175 billion parameters but was trained on only 300 billion tokens, less than a tenth of what Chinchilla recommends. Later models shifted toward more data and not-too-large models. For the derivation and how to extrapolate from small experiments, CS336 Lecture 9 goes further.
What pretraining costs
The slides sort the challenges into two groups:
- Cost: at least millions of dollars, lots of time, lots of electricity
- Learned knowledge: there's a "knowledge cutoff" (the slides include a screenshot of the cutoff date on an OpenAI model page). Learned knowledge is hard to edit. And there's what the slides put in quotes as "plagiarism," meaning the model may reproduce its training data
The pretraining bill: memory runs out first
Why is pretraining so expensive? The slides start by listing what one training step has to store:
| Stage | What gets stored | Size depends on |
|---|---|---|
| Initialization | Model parameters | Billions to hundreds of billions |
| Forward pass | Activations (needed to compute the loss) | Model size, batch size, context length |
| Backward pass | Gradients | Same count as the parameters |
| Weight update | Optimizer state | Adam keeps two extra values per parameter |
Formula: why Adam eats memory
θ_{t+1} ← θ_t − α · m_t / (√v_t + ε)
m_{t+1} ← β₁ m_t + (1 − β₁) ∇L(θ_t) # moving average of the gradient
v_{t+1} ← β₂ v_t + (1 − β₂) (∇L(θ_t))² # moving average of the squared gradient
m and v are each as large as the parameters, so the optimizer state alone is twice the parameter count. The slides suggest the original Adam and AdamW papers as further reading.
The problem is that a single GPU has only tens of GB of memory. The slides use the NVIDIA H100 spec sheet as the example, boxing the GPU memory row (80GB and 94GB variants). A model with hundreds of billions of parameters can't even fit its parameters, let alone gradients and optimizer state.
Saving memory and time: a map
The lecture then spends roughly half its slides on training optimizations. In the 2026 edition these topics were pulled out into their own lecture, "LLM systems" (covered in order 10 of this series), and the site already has in-depth CS336 posts on them, so here you get only the map, with each cell stating what problem it solves:
| Technique | Problem it solves | What the slides emphasize | Go deeper |
|---|---|---|---|
| Data parallelism | Too much data for one device | Split the batch across GPUs; each holds the full model | CS336 Lecture 7 |
| ZeRO | Every GPU stores the same things, which is wasteful | ZeRO-1 shards optimizer state; ZeRO-2 adds gradients; ZeRO-3 adds parameters | CS336 Lecture 8 |
| Model parallelism | The model itself doesn't fit on one GPU | Split computation across devices: tensor (TP), pipeline (PP), sequence (SP), context (CP), expert (EP) parallelism | The slides recommend Hugging Face's Ultra-Scale Playbook |
| FlashAttention | Attention keeps shuttling data between slow HBM and fast SRAM | Tile blocks into SRAM, compute, then write back; in the backward pass, recompute rather than store the big matrix; the result is exact, not an approximation | CS336 Lecture 5 |
| Mixed precision training | FP32 is slow and takes space | Activations in the forward pass and gradients in the backward pass in low precision; weight updates keep high precision | CS336 Lecture 2 |
The FlashAttention row is worth a second look because it's counterintuitive. The slides quote a set of numbers from the original paper: FlashAttention actually does more computation (75.2 vs 66.6 GFLOPs), but HBM reads/writes drop from 40.3 GB to 4.4 GB and runtime drops from 41.7 ms to 7.3 ms. On a GPU the bottleneck is often moving data, not computing. In the slides' own words: "More FLOPs, but less runtime!!"
The mixed-precision row requires knowing how floats are laid out. The slides list the bit split for each format:
| Format | Exponent bits | Mantissa bits |
|---|---|---|
| FP32 | 8 | 23 |
| FP16 | 5 | 10 |
| BF16 | 8 | 7 |
BF16 keeps the same exponent range as FP32 and gives up precision. The slides' takeaway is that lower precision means faster processing on the GPU.
Inference-side optimizations (KV cache, speculative decoding, and so on) came up in earlier lectures and are outside this one; for those, see CS336 Lecture 10.
Stage two: SFT, or "graduating" the model into an assistant
How it works
SFT (Supervised FineTuning) changes model behavior by tuning its weights:
- Collect input/desired-output pairs, the SFT data
- Train with the same next-token objective, except now conditioned on the input, so the model learns only how to produce the output
When the data is instruction/response pairs, this is called instruction tuning, and the slides cite the FLAN paper. Examples include a short story about a teddy bear who likes poetry, three things a teddy bear might do on a rainy day, a poem, and an explanation of why a teddy bear is a great friend.
Data
The data can be human-written or synthetic. The slides list four kinds: assistant dialogs, synthetic instructions, math/reasoning/code, and safety alignment. The scale runs from thousands to millions of examples:
| Model | SFT examples (per the slides' table) |
|---|---|
| GPT-3 family | 13 thousand |
| Llama 3 | 10 million |
Compared with trillions of pretraining tokens, SFT is several orders of magnitude smaller, which is why many more teams can afford it.
What makes it hard
The slides list five challenges: it needs very high-quality data, it's sensitive to the prompt distribution, generalization, it's hard to evaluate, and it's still computationally expensive.
The slides expand on evaluation along two lines:
- Benchmarks: MMLU for general knowledge, ARC-Challenge for basic reasoning, GSM8K for math, HumanEval for code. The slides note that to compare across models, it's recommended to train on the test task first, citing Dominguez-Olmedo et al.: whether a model was trained on the test task confounds evaluation
- "Real-life" feel: sites like Chatbot Arena (now LMArena) let users vote in A/B tests between two anonymous models' answers. The upside is that it puts a number on "vibes." The problems include a cold start from unequal exposure of models, being easy to rig, users being unable to judge things like factuality, non-representative personal preferences, and penalizing safety refusals
The slides conclude: "evaluation is a hard problem in itself!" Order 8 of this series, on evaluation, is devoted to it.
There's one more step after SFT. The final lifecycle diagram in the slides reads: initialization → pretraining → SFT → preference tuning, which makes the model "not misbehave as much"; together these are called alignment. That's the topic of Lecture 5.
LoRA: finetuning without big GPUs
Intuition
SFT updates all of the model's weights, and the slides say it plainly: "not everyone has big GPUs." LoRA (Low-Rank Adaptation) freezes the original weight matrix W₀ and trains only a correction, expressed as the product of two thin matrices.
The slides mention three benefits:
- Only a small fraction of parameters are trained, with performance similar to full finetuning
- Swapping matrices means swapping tasks: keep one W₀, attach the spam-detection B and A to do spam detection, swap in the sentiment B and A to do sentiment extraction, with no need to store several full models
- Related methods include prefix tuning and adapters
Formula: how much the low-rank decomposition saves
W = W₀ + B · A
W₀ : d × k, frozen
B : d × r, trained
A : r × k, trained
r ≪ min(d, k)
A worked example (these numbers are not from the slides): a 4096 × 4096 matrix has about 16.77 million parameters. With r = 8, B and A together have only 4096×8 + 8×4096 = 65,536, about 0.4%.
Which layers to apply it to
This is one of the newer parts of the deck. The original LoRA paper experimented only on attention weight matrices; the slides cite Thinking Machines' 2025 post LoRA Without Regret for today's guidance: apply it to both attention and the feed-forward layers (FFN), with the feed-forward layer as the most important location.
The same source is cited for two training differences, which the slides label as empirical:
- LoRA needs a higher learning rate than full finetuning
- LoRA does worse than full finetuning at large batch sizes (the post says "in some scenarios," and raising the rank doesn't fix it)
The site has a separate write-up of LoRA Without Regret: CS224N Lecture 18 materials notes.
QLoRA: shrinking the frozen weights too
LoRA saves on gradients and optimizer state, but the frozen W₀ still has to sit in memory in full. QLoRA quantizes the frozen weights for storage:
- Storage: frozen weights in 4-bit; LoRA's B and A in full precision
- Computation: dequantized to higher precision when used
- NF4: ordinary INT8 quantization splits the value range uniformly; NF4 (4-bit NormalFloat) splits by the quantiles of a normal distribution, which wastes less because neural network weights are roughly normally distributed
- Double quantization: quantization needs stored "quantization constants," and QLoRA quantizes those constants a second time
The slides cite the paper's LLaMA 65B results: about 16x VRAM savings during finetuning, with double quantization saving an extra ~6%.
Connecting back to the models you use
The model you use in a chat interface today has almost certainly gone through every box in that diagram: pretraining, SFT, then preference tuning. When it tells you its knowledge stops at a certain date, that's the edge of its pretraining data. When it follows your requested format, that's SFT at work.
If you're finetuning yourself, the lecture has three practical takeaways:
- Decide whether you're changing behavior or knowledge. SFT is good at changing behavior (format, tone, following instructions); the slides list "hard to edit knowledge" as a pretraining challenge.
- Update your LoRA defaults. If your config attaches LoRA only to attention, follow LoRA Without Regret: add the FFN layers and re-sweep with a higher learning rate.
- Evaluation is harder than training. Winning on benchmarks doesn't mean the model feels good in real use, and Arena scores have their own biases. Look at both.
What changed in 2026
Apart from Lecture 1, the 2026 slides haven't been released yet, so the following compares only the topic lists on the 2026 syllabus:
- The training lecture grows into a full post-training pipeline: 2026 Lecture 3, "LLM training" (October 9), lists pretraining, SFT, and LoRA, then pulls in preference tuning (RLHF, DPO) and reasoning, which 2025 covered in Lectures 5 and 6, and adds on-policy distillation and "distillation to smaller models."
- Systems optimization becomes its own lecture: 2026 Lecture 5, "LLM systems" (October 30), lists distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, Flash Attention, and hardware trade-offs. This lecture's data parallelism, ZeRO, and FlashAttention will likely move there; see order 10 of this series.
- Quantization doesn't appear on the 2026 syllabus: neither lecture's topic list names quantization, mixed precision, or QLoRA. They may be folded into "hardware trade-offs" or the LoRA section; that can only be confirmed once the slides are out.
Self-check
These questions are adapted from Part IV, "LLM training," of the 2025 midterm; answers are in the solutions PDF:
- What best describes SFT? How does it differ from training a reward model or running PPO? (Q3)
- In mixed-precision training as practiced, what is kept in low precision and what in high precision? (Q5)
- What does FlashAttention mainly optimize? Is it exact or approximate? (Q6)
- Which ZeRO variant shards optimizer state, gradients, and parameters across devices? (Q7)
- In QLoRA, what format stores the frozen weights? What precision do the LoRA matrices and matrix multiplications use? (Q8)
- Compared with a pretrained model, what does instruction tuning try to achieve? Name two practical challenges discussed in lecture. (Q9)
Going deeper
- Another take on pretraining: CS224N Lecture 7: Pretraining, subwords, and in-context learning
- LoRA and other parameter-efficient methods: CS224N Lecture 9: Prompting, LoRA, and parameter-efficient finetuning
- Counting FLOPs and memory yourself: CS336 Lecture 2
- Why GPUs are bottlenecked on data movement: CS336 Lecture 5
- ZeRO, FSDP, and 3D parallelism: CS336 Lecture 8
- RLHF after SFT: CS336 Lecture 15, and Lecture 5 of this series
References
- CME 295 2025 syllabus
- CME 295 2026 syllabus
- 2025 Lecture 4 slides (PDF)
- 2025 Lecture 4 recording
- 2025 midterm / solutions
- Brown et al., Language Models are Few-Shot Learners (2020)
- Llama Team, The Llama 3 Herd of Models (2024)
- Kaplan et al., Scaling Laws for Neural Language Models (2020)
- Hoffmann et al., Training Compute-Optimal Large Language Models (2022)
- Kingma & Ba, Adam (2014)
- Loshchilov & Hutter, Decoupled Weight Decay Regularization (2017)
- NVIDIA H100 Tensor Core GPU
- Rajbhandari et al., ZeRO (2019)
- Hugging Face, The Ultra-Scale Playbook (2025)
- Dao et al., FlashAttention (2022)
- Micikevicius et al., Mixed Precision Training (2017)
- Wei et al., Finetuned Language Models Are Zero-Shot Learners (2021)
- Dominguez-Olmedo et al., Training on the Test Task Confounds Evaluation and Emergence (2024)
- LMArena Leaderboard
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021)
- Schulman et al., LoRA Without Regret (Thinking Machines, 2025)
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (2023)
- Reading Stanford CME295 (series overview)
Loading...