Skip to content

CME295 Lecture 4: The Bill for Training an LLM, and Where Pretraining, SFT, and LoRA Spend It

Sep 29, 20261 min
TL;DRCME295 Lecture 4 splits LLM training into two stages: pretraining on trillions of tokens (Llama 3 used 15 trillion), then SFT on thousands to millions of demonstrations so the model stops continuing text and starts answering. In between sits a map of memory savers (ZeRO, FlashAttention, mixed precision); the lecture closes with LoRA and QLoRA, which let people without big GPUs finetune, with QLoRA cutting VRAM by about 16x on a 65B model.

🌏 中文版

This post covers Lecture 4, "LLM training," of the 2025 edition of Stanford's CME295 (October 17, 2025). The main source is the 128-page slide deck; the recording is there to watch alongside. Everything below is based only on the text and figures on the slides, not on anything said out loud in class.

The first three lectures answered "what does an LLM look like?" Lecture 4 asks a different question: how does that machine go from random weights to an assistant that answers questions? The slides give a two-part answer. First pretraining, which teaches the model the patterns of language and code. Then finetuning, which teaches it to do what it's told. The two stages differ in cost by several orders of magnitude, and most of the lecture is about where the money and memory go and how to save them.

Start with a counterexample: a pretrained model only continues text

The slides open with a question: "Can I put my teddy bear in the washer?"

A model that has only been pretrained replies: "Teddy bears are often made of materials like polyester and cotton, with plastic eyes and sometimes small accessories." Nothing in it is wrong. It is simply continuing with the kind of text a web page might contain, and never answers the question.

A model that has been instruction tuned replies: "No, it might get damaged. Try hand washing instead."

The gap between those two answers is the two stages of this lecture.

flowchart LR
  I["Randomly initialized model"] -->|"Pretraining<br/>trillions of tokens, next-token prediction<br/>cost: millions of dollars at least"| P["Pretrained model<br/>knows language and code, but only continues text"]
  P -->|"SFT / instruction tuning<br/>thousands to millions of demonstrations"| S["Assistant model<br/>follows instructions"]
  P -.->|"LoRA / QLoRA<br/>train only a small fraction of parameters"| S
  S -->|"Preference tuning (Lecture 5)"| A["Model that misbehaves less"]

The two-stage recipe is itself a paradigm shift. The slides contrast three approaches. Traditional machine learning trains a model from scratch for every task: one for spam detection, one for sentiment extraction, one for translation. Transfer learning reuses part of a trained model. LLM training first trains a model to understand language, then tunes it for the end task.

Stage one: pretraining

Objective and data

The pretraining objective is simple: predict the next token. Given "[BOS] A teddy bear is", guess the next word.

The goal is to learn "patterns of language and code," so the data mixes web-scraped text (the slides name Common Crawl and Wikipedia) with code (GitHub, Stack Overflow), across many languages. The scale is trillions of tokens:

ModelPretraining tokens
GPT-3 (2020)300 billion
Llama 3 (2024)15 trillion

Measuring scale in FLOPs

The slides separate two terms that often get mixed up. FLOPs is the total amount of computation (how many floating-point operations were done); FLOPS or FLOP/s is how many can be done per second. The first is the training bill, the second is hardware speed. In orders of magnitude, training a small neural network takes about 10⁷ FLOPs, a large RNN about 10¹⁴, an LLM about 10²⁵. On the hardware side, a smartphone does about 10¹² FLOPS, a computer about 10¹⁴, a supercomputer about 10¹⁸.

Scaling laws: bigger models are more sample-efficient

The slides take two findings from Kaplan et al.'s scaling-law paper:

  • Scaling: test loss falls as a power law in compute, dataset size, and parameter count, each a straight line on log-log axes
  • Sample efficiency: larger models reach lower loss with the same number of tokens, so bigger models learn faster

Then comes the Chinchilla law. The slides reproduce the paper's table of "compute-optimal" training tokens for different model sizes:

ParametersOptimal training tokens
1 billion20.2 billion
67 billion1.5 trillion
175 billion3.7 trillion

Reading off the table, that's roughly 20 tokens per parameter. Put it next to the earlier table: GPT-3 has 175 billion parameters but was trained on only 300 billion tokens, less than a tenth of what Chinchilla recommends. Later models shifted toward more data and not-too-large models. For the derivation and how to extrapolate from small experiments, CS336 Lecture 9 goes further.

What pretraining costs

The slides sort the challenges into two groups:

  • Cost: at least millions of dollars, lots of time, lots of electricity
  • Learned knowledge: there's a "knowledge cutoff" (the slides include a screenshot of the cutoff date on an OpenAI model page). Learned knowledge is hard to edit. And there's what the slides put in quotes as "plagiarism," meaning the model may reproduce its training data

The pretraining bill: memory runs out first

Why is pretraining so expensive? The slides start by listing what one training step has to store:

StageWhat gets storedSize depends on
InitializationModel parametersBillions to hundreds of billions
Forward passActivations (needed to compute the loss)Model size, batch size, context length
Backward passGradientsSame count as the parameters
Weight updateOptimizer stateAdam keeps two extra values per parameter
Formula: why Adam eats memory
θ_{t+1} ← θ_t − α · m_t / (√v_t + ε)

m_{t+1} ← β₁ m_t + (1 − β₁) ∇L(θ_t)        # moving average of the gradient
v_{t+1} ← β₂ v_t + (1 − β₂) (∇L(θ_t))²     # moving average of the squared gradient

m and v are each as large as the parameters, so the optimizer state alone is twice the parameter count. The slides suggest the original Adam and AdamW papers as further reading.

The problem is that a single GPU has only tens of GB of memory. The slides use the NVIDIA H100 spec sheet as the example, boxing the GPU memory row (80GB and 94GB variants). A model with hundreds of billions of parameters can't even fit its parameters, let alone gradients and optimizer state.

Saving memory and time: a map

The lecture then spends roughly half its slides on training optimizations. In the 2026 edition these topics were pulled out into their own lecture, "LLM systems" (covered in order 10 of this series), and the site already has in-depth CS336 posts on them, so here you get only the map, with each cell stating what problem it solves:

TechniqueProblem it solvesWhat the slides emphasizeGo deeper
Data parallelismToo much data for one deviceSplit the batch across GPUs; each holds the full modelCS336 Lecture 7
ZeROEvery GPU stores the same things, which is wastefulZeRO-1 shards optimizer state; ZeRO-2 adds gradients; ZeRO-3 adds parametersCS336 Lecture 8
Model parallelismThe model itself doesn't fit on one GPUSplit computation across devices: tensor (TP), pipeline (PP), sequence (SP), context (CP), expert (EP) parallelismThe slides recommend Hugging Face's Ultra-Scale Playbook
FlashAttentionAttention keeps shuttling data between slow HBM and fast SRAMTile blocks into SRAM, compute, then write back; in the backward pass, recompute rather than store the big matrix; the result is exact, not an approximationCS336 Lecture 5
Mixed precision trainingFP32 is slow and takes spaceActivations in the forward pass and gradients in the backward pass in low precision; weight updates keep high precisionCS336 Lecture 2

The FlashAttention row is worth a second look because it's counterintuitive. The slides quote a set of numbers from the original paper: FlashAttention actually does more computation (75.2 vs 66.6 GFLOPs), but HBM reads/writes drop from 40.3 GB to 4.4 GB and runtime drops from 41.7 ms to 7.3 ms. On a GPU the bottleneck is often moving data, not computing. In the slides' own words: "More FLOPs, but less runtime!!"

The mixed-precision row requires knowing how floats are laid out. The slides list the bit split for each format:

FormatExponent bitsMantissa bits
FP32823
FP16510
BF1687

BF16 keeps the same exponent range as FP32 and gives up precision. The slides' takeaway is that lower precision means faster processing on the GPU.

Inference-side optimizations (KV cache, speculative decoding, and so on) came up in earlier lectures and are outside this one; for those, see CS336 Lecture 10.

Stage two: SFT, or "graduating" the model into an assistant

How it works

SFT (Supervised FineTuning) changes model behavior by tuning its weights:

  1. Collect input/desired-output pairs, the SFT data
  2. Train with the same next-token objective, except now conditioned on the input, so the model learns only how to produce the output

When the data is instruction/response pairs, this is called instruction tuning, and the slides cite the FLAN paper. Examples include a short story about a teddy bear who likes poetry, three things a teddy bear might do on a rainy day, a poem, and an explanation of why a teddy bear is a great friend.

Data

The data can be human-written or synthetic. The slides list four kinds: assistant dialogs, synthetic instructions, math/reasoning/code, and safety alignment. The scale runs from thousands to millions of examples:

ModelSFT examples (per the slides' table)
GPT-3 family13 thousand
Llama 310 million

Compared with trillions of pretraining tokens, SFT is several orders of magnitude smaller, which is why many more teams can afford it.

What makes it hard

The slides list five challenges: it needs very high-quality data, it's sensitive to the prompt distribution, generalization, it's hard to evaluate, and it's still computationally expensive.

The slides expand on evaluation along two lines:

  • Benchmarks: MMLU for general knowledge, ARC-Challenge for basic reasoning, GSM8K for math, HumanEval for code. The slides note that to compare across models, it's recommended to train on the test task first, citing Dominguez-Olmedo et al.: whether a model was trained on the test task confounds evaluation
  • "Real-life" feel: sites like Chatbot Arena (now LMArena) let users vote in A/B tests between two anonymous models' answers. The upside is that it puts a number on "vibes." The problems include a cold start from unequal exposure of models, being easy to rig, users being unable to judge things like factuality, non-representative personal preferences, and penalizing safety refusals

The slides conclude: "evaluation is a hard problem in itself!" Order 8 of this series, on evaluation, is devoted to it.

There's one more step after SFT. The final lifecycle diagram in the slides reads: initialization → pretraining → SFT → preference tuning, which makes the model "not misbehave as much"; together these are called alignment. That's the topic of Lecture 5.

LoRA: finetuning without big GPUs

Intuition

SFT updates all of the model's weights, and the slides say it plainly: "not everyone has big GPUs." LoRA (Low-Rank Adaptation) freezes the original weight matrix W₀ and trains only a correction, expressed as the product of two thin matrices.

The slides mention three benefits:

  • Only a small fraction of parameters are trained, with performance similar to full finetuning
  • Swapping matrices means swapping tasks: keep one W₀, attach the spam-detection B and A to do spam detection, swap in the sentiment B and A to do sentiment extraction, with no need to store several full models
  • Related methods include prefix tuning and adapters
Formula: how much the low-rank decomposition saves
W = W₀ + B · A

W₀ : d × k, frozen
B  : d × r, trained
A  : r × k, trained
r  ≪ min(d, k)

A worked example (these numbers are not from the slides): a 4096 × 4096 matrix has about 16.77 million parameters. With r = 8, B and A together have only 4096×8 + 8×4096 = 65,536, about 0.4%.

Which layers to apply it to

This is one of the newer parts of the deck. The original LoRA paper experimented only on attention weight matrices; the slides cite Thinking Machines' 2025 post LoRA Without Regret for today's guidance: apply it to both attention and the feed-forward layers (FFN), with the feed-forward layer as the most important location.

The same source is cited for two training differences, which the slides label as empirical:

  • LoRA needs a higher learning rate than full finetuning
  • LoRA does worse than full finetuning at large batch sizes (the post says "in some scenarios," and raising the rank doesn't fix it)

The site has a separate write-up of LoRA Without Regret: CS224N Lecture 18 materials notes.

QLoRA: shrinking the frozen weights too

LoRA saves on gradients and optimizer state, but the frozen W₀ still has to sit in memory in full. QLoRA quantizes the frozen weights for storage:

  • Storage: frozen weights in 4-bit; LoRA's B and A in full precision
  • Computation: dequantized to higher precision when used
  • NF4: ordinary INT8 quantization splits the value range uniformly; NF4 (4-bit NormalFloat) splits by the quantiles of a normal distribution, which wastes less because neural network weights are roughly normally distributed
  • Double quantization: quantization needs stored "quantization constants," and QLoRA quantizes those constants a second time

The slides cite the paper's LLaMA 65B results: about 16x VRAM savings during finetuning, with double quantization saving an extra ~6%.

Connecting back to the models you use

The model you use in a chat interface today has almost certainly gone through every box in that diagram: pretraining, SFT, then preference tuning. When it tells you its knowledge stops at a certain date, that's the edge of its pretraining data. When it follows your requested format, that's SFT at work.

If you're finetuning yourself, the lecture has three practical takeaways:

  1. Decide whether you're changing behavior or knowledge. SFT is good at changing behavior (format, tone, following instructions); the slides list "hard to edit knowledge" as a pretraining challenge.
  2. Update your LoRA defaults. If your config attaches LoRA only to attention, follow LoRA Without Regret: add the FFN layers and re-sweep with a higher learning rate.
  3. Evaluation is harder than training. Winning on benchmarks doesn't mean the model feels good in real use, and Arena scores have their own biases. Look at both.

What changed in 2026

Apart from Lecture 1, the 2026 slides haven't been released yet, so the following compares only the topic lists on the 2026 syllabus:

  • The training lecture grows into a full post-training pipeline: 2026 Lecture 3, "LLM training" (October 9), lists pretraining, SFT, and LoRA, then pulls in preference tuning (RLHF, DPO) and reasoning, which 2025 covered in Lectures 5 and 6, and adds on-policy distillation and "distillation to smaller models."
  • Systems optimization becomes its own lecture: 2026 Lecture 5, "LLM systems" (October 30), lists distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, Flash Attention, and hardware trade-offs. This lecture's data parallelism, ZeRO, and FlashAttention will likely move there; see order 10 of this series.
  • Quantization doesn't appear on the 2026 syllabus: neither lecture's topic list names quantization, mixed precision, or QLoRA. They may be folded into "hardware trade-offs" or the LoRA section; that can only be confirmed once the slides are out.

Self-check

These questions are adapted from Part IV, "LLM training," of the 2025 midterm; answers are in the solutions PDF:

  1. What best describes SFT? How does it differ from training a reward model or running PPO? (Q3)
  2. In mixed-precision training as practiced, what is kept in low precision and what in high precision? (Q5)
  3. What does FlashAttention mainly optimize? Is it exact or approximate? (Q6)
  4. Which ZeRO variant shards optimizer state, gradients, and parameters across devices? (Q7)
  5. In QLoRA, what format stores the frozen weights? What precision do the LoRA matrices and matrix multiplications use? (Q8)
  6. Compared with a pretrained model, what does instruction tuning try to achieve? Name two practical challenges discussed in lecture. (Q9)

Going deeper

References