🌏 中文版
Version note: This post follows the Spring 2026 offering of CMU 11-868 LLM Systems. The main source is the April 8 Lecture 23 slides, Parameter Efficient Fine-Tuning for LLM (Lei Li, 46 PDF pages; page numbers below are PDF page order, not the numbers printed in slide corners), plus the three readings the Syllabus lists: CIAT, LoRA, and QLoRA. All facts were checked against the official materials on 2026-09-30. Access level A3: the slides are public. What you can't get is lecture video (this course publishes none) and the Canvas quiz.
Series: Previous L12–L13 TPU, JAX, and Pallas/Splash Attention | Next L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention | Series overview
The earlier lectures kept asking how to make one training run faster or make it fit at all. Lecture 23 changes the setting. The model is already pretrained. You want it to learn a new task or domain, and you have one or two GPUs.
The lecture answers in two layers. First, LoRA: freeze the original weights and train only a low-rank update. Second, QLoRA: compress even the frozen weights to 4 bits. Together, the slides take LLaMA-8B fine-tuning memory from about 80GB to about 9.2GB.
Page 2 recaps the previous lecture on quantization (absmax, zero-point, LLM.int8(), GPTQ), because QLoRA builds on it. If you skipped it, read L19–L20 Model Quantization first.
The bill: why full fine-tuning is expensive
Pages 4–6 compare three approaches with one "bits per parameter" table:
| Approach | Weights | Weight gradients | Optimizer state | Adapter weights | LLaMA-8B working memory |
|---|---|---|---|---|---|
| Half-precision full fine-tuning (Adam) | 16 bit | 16 bit | 2 × 16 bit | — | ~80GB |
| LoRA | 16 bit | ~0.4 bit | ~0.8 bit | ~0.4 bit | ~33GB, fits one A100 |
| QLoRA | 4 bit | ~0.4 bit | ~0.8 bit | ~0.4 bit | ~9.2GB |
All three rows list activations as roughly 1–2× the parameter count, independent of method. The way to read the table: the expensive part of full fine-tuning isn't the weights. It's the gradients plus Adam's two moments, and those are exactly what LoRA cuts.
Page 7 sorts PEFT into three families, following the survey by Lialin et al.:
- Selective: fine-tune a chosen subset of parameters.
- Reparameterization: represent the weight update in low rank. LoRA lives here.
- Additive: add new trainable layers or parameters, such as adapters or soft prompts.
LoRA: freeze W, learn A·B
The formula on pages 9–10 is one line. Freeze the pretrained weight W and learn a low-rank increment:
W′ = W + A·B, where A is d×r, B is r×d, and r is much smaller than d.
Page 10 also records where the idea came from. It first appeared in the multilingual translation paper Counter-Interference Adapter (CIAT, EMNLP 2021), and was later "re-invented" by LoRA (ICLR 2022). Their first arXiv versions are from April and June 2021. From here on the slides write "LoRA/CIAT".
Which matrices get it. Page 11 recommends rank 8 or 16, applied to attention's W_Q, W_K, W_V, and W_O, and not to the FFN's linear layers. The slide's reason is that the FFN is "storing knowledge". The same page adds that CIAT also applies it to embeddings and FFN layers, and that this helps. Read together, "attention only" is a default, not a law.
At inference (page 12). Compute W′ = W + A·B and use it as an ordinary weight; the only extra storage is A and B. To switch tasks, subtract the old adapter and add the new one: W″ = W′ − A·B + A″·B″. One base model can carry many small domain-specific adapters.
Backprop needs only two small gradients (page 13). W₀ is fixed, so you differentiate only with respect to A and B:
- ∂L/∂A = g_out · (Bx)ᵀ
- ∂L/∂B = (g_out · A) · xᵀ (as written on the slide; with the transposes spelled out by dimension it is (Aᵀ · g_out) · xᵀ)
If you did HW2 MiniTorch in this series, this is the matmul backward rule applied twice.
What you store (page 14): the original parameters, plus adapter weights, adapter gradients, and the adapter's two Adam moments (each 2×d×r), plus activations. The original parameters need no gradients or optimizer state.
Page 15 puts numbers side by side. LLaMA 8B's 8 billion parameters take 16GB in BF16, and Adam's optimizer state takes 48GB. With LoRA, trainable parameters drop to about 4M, which means 8MB of weights and 24MB of optimizer state. Page 16 concludes that memory and training time both go down, and inference works like any other LLM.
How large should the rank be
Pages 17–21 summarize the LoRA paper's experiments. NLU uses RoBERTa and DeBERTa on eight GLUE subtasks; NLG uses GPT-2 and GPT-3. Baselines are full fine-tuning, BitFit, two kinds of prefix tuning, and adapter tuning. The slides pull out two points:
- Page 19: raising the rank doesn't cover more meaningful subspaces. A low-rank matrix with r=8 is enough.
- Page 20: in a stress test scaled to 175B-parameter GPT-3, not every method improves monotonically as trainable parameters grow.
Page 21 also shows a chart of LoRA on GSM8K math problems. The slide gives no text stating a conclusion, so this post doesn't draw one for it.
QLoRA: put the frozen weights in 4 bits too
Page 23 separates two kinds of quantization. The previous lecture's GPTQ is post-training quantization: convert a trained model to lower precision, with no retraining. QLoRA is filed under quantization-aware training: quantization happens during training, which usually performs better.
Page 24 lists the three innovations of QLoRA (NeurIPS 2023): 4-bit NormalFloat storage, double quantization, and paged optimizers, with all computation in BF16. The slide cites the paper's result: average memory for fine-tuning a 65B model drops from 780GB to 48GB on a single GPU, with performance on par with 16-bit fine-tuning.
The algorithm on page 25 is a "store in 4 bits, compute in 16 bits" pipeline:
- Store the weights in NF4 with double quantization.
- Dequantize to BF16 when they're needed.
- Run forward and backward in BF16.
- Compute BF16 weight gradients only for the LoRA parameters.
NF4: why not equal-width bins
Four bits give you 16 values. Page 26 points out that absmax-style equal-width bins do badly on unevenly distributed data: lots of small values that differ only slightly get quantized into the same bin.
NF4 on page 27:
- Group weights in blocks of 64.
- Take the block's max absolute value as the scale.
- Divide the 64 numbers by the scale so they fall in [−1, 1].
- Find the nearest value in a 16-entry lookup table and output its index (−7 to 8).
The table isn't evenly spaced: bins are dense near 0 and sparse near ±1. Page 29 says the values come from quantiles ("probably quantile").
Pages 28 and 30 work one example by hand. It's worth following along:
- Original value 0.0045, block scale 0.0071, normalized to about 0.63.
- The nearest table entry is 0.5626, index 6.
- Dequantized: 0.5626 × 0.0071 ≈ 0.00399, an error of about 0.0005.
Double quantization: the scales need compressing too
Page 31 names the cost of small blocks. One 32-bit constant per 64 parameters adds 32/64 = 0.5 bit per parameter on average.
Page 32 quantizes again. Each block of 64 values gets NF4; the resulting scales are grouped 256 at a time and quantized to FP8; only the second-level scale is stored in FP32. The slide's total is about 0.52 bytes per parameter.
Page 33 collects the full QLoRA configuration: pretrained weights in NF4 with double quantization, LoRA weights in BF16, and inputs plus all computation (forward, backward, optimizer) in BF16.
Paged optimizers: spill to CPU before you crash
Pages 34–36 deal with occasional OOMs during training. Optimizer state is allocated in NVIDIA unified memory. When the GPU runs short, those pages move to the CPU automatically, and they come back to the GPU when the optimizer update step needs them. The slides compare it to ordinary paging between CPU RAM and disk.
Pages 37–38 cite two results from the paper. LoRA's default hyperparameters don't match 16-bit performance, and NF4 beats 4-bit float (FP4). Once tuned, 4-bit QLoRA matches both 16-bit full fine-tuning and 16-bit LoRA.
Code walkthrough
Pages 40–42 are screenshots of a LoRA layer implementation, with three points: initialize the A and B layers; freeze the pretrained weight; in the forward pass, sum the outputs of the two branches (original weight and low-rank branch). After training, the LoRA weights can be merged back into the original weight.
Page 43 shows QLoRA usage: load the model with a quantization config, then apply LoRA as usual. The CUDA quantization ops come from wrappers in bitsandbytes. The slide's configuration looks like this (model name as on the slide):
AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-70b",
quantization_config=BitsAndBytesConfig(
load_in_4bit=args.bits == 4,
load_in_8bit=args.bits == 8,
llm_int8_threshold=6.0,
llm_int8_has_fp16_weight=False,
bnb_4bit_compute_dtype=compute_dtype,
bnb_4bit_use_double_quant=args.double_quant,
bnb_4bit_quant_type=args.quant_type,
),
)
The same page links a Guanaco 7B Colab example in the QLoRA repo.
How this connects to the homework
LoRA shows up directly in HW6. The assignment has you turn on LoRA on top of DeepSpeed ZeRO so Llama-2-7B training fits on two 16GB V100s. The memory table from this lecture is the one you'll be doing in your head while tuning that run. For the ZeRO half, read L18 ZeRO first.
The summary on pages 44–45: LoRA/CIAT adapts large models cheaply with a small set of low-rank parameters, works at GPT-3 scale, and does better in low-data settings. QLoRA stores weights with NF4 double quantization and computes in BF16, matching the original LoRA. The slides call it the first method that can fine-tune a 33B LLM on a single consumer GPU.
How to study it yourself
- Read pages 4–6 and recompute the three rows yourself as "bits per parameter × parameter count". Make sure you know which part got cut.
- Follow pages 28 and 30 to quantize and dequantize with NF4 by hand, then repeat with a different number.
- Read the rank experiments in the LoRA paper and compare them with page 19.
- If you have a GPU, load a small model with the page 43 configuration and compare memory with
bnb_4bit_use_double_quanton and off.
Further reading
- Where the quantization comes from: L19–L20 Model Quantization
- Another way to save memory on the training side: L18 ZeRO
- Estimating training cost from memory and compute: CS336 Resource Accounting
References
- CMU 11-868 LLM Systems (Spring 2026) course home
- 11-868 Syllabus (Spring 2026) — the April 8 session and its CIAT, LoRA, and QLoRA readings
- Lecture 23 slides: Parameter Efficient Fine-Tuning for LLM — memory math, LoRA/CIAT, NF4, double quantization, paged optimizers, code walkthrough
- Zhu et al., Counter-Interference Adapter for Multilingual Machine Translation (arXiv 2104.08154)
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (arXiv 2106.09685)
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (arXiv 2305.14314)
- Lialin et al., Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning (arXiv 2303.15647) — source of the PEFT taxonomy on page 7
- artidoro/qlora GitHub repo — the Guanaco example linked on page 43
- bitsandbytes — the quantization wrappers used in the QLoRA code
Loading...