Skip to content

CMU 11-868 L23 Efficient Fine-Tuning for Large Models: LoRA, CIAT, and QLoRA

Sep 30, 20261 min
TL;DRLecture 23 of 11-868 treats fine-tuning as a memory problem. Full-parameter half-precision fine-tuning of LLaMA-8B needs about 80GB. LoRA brings that to about 33GB, and QLoRA, which stores the frozen weights in 4 bits, gets it to about 9.2GB. The lecture moves in three steps: train only two small low-rank matrices A and B (the slides credit CIAT as the first to do this); squeeze frozen weights to about 0.52 bytes per parameter with an NF4 lookup table and double quantization; and use a paged optimizer to push optimizer state to the CPU when the GPU is about to run out.

🌏 中文版

Version note: This post follows the Spring 2026 offering of CMU 11-868 LLM Systems. The main source is the April 8 Lecture 23 slides, Parameter Efficient Fine-Tuning for LLM (Lei Li, 46 PDF pages; page numbers below are PDF page order, not the numbers printed in slide corners), plus the three readings the Syllabus lists: CIAT, LoRA, and QLoRA. All facts were checked against the official materials on 2026-09-30. Access level A3: the slides are public. What you can't get is lecture video (this course publishes none) and the Canvas quiz.

Series: Previous L12–L13 TPU, JAX, and Pallas/Splash Attention | Next L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention | Series overview

The earlier lectures kept asking how to make one training run faster or make it fit at all. Lecture 23 changes the setting. The model is already pretrained. You want it to learn a new task or domain, and you have one or two GPUs.

The lecture answers in two layers. First, LoRA: freeze the original weights and train only a low-rank update. Second, QLoRA: compress even the frozen weights to 4 bits. Together, the slides take LLaMA-8B fine-tuning memory from about 80GB to about 9.2GB.

Page 2 recaps the previous lecture on quantization (absmax, zero-point, LLM.int8(), GPTQ), because QLoRA builds on it. If you skipped it, read L19–L20 Model Quantization first.

The bill: why full fine-tuning is expensive

Pages 4–6 compare three approaches with one "bits per parameter" table:

ApproachWeightsWeight gradientsOptimizer stateAdapter weightsLLaMA-8B working memory
Half-precision full fine-tuning (Adam)16 bit16 bit2 × 16 bit—~80GB
LoRA16 bit~0.4 bit~0.8 bit~0.4 bit~33GB, fits one A100
QLoRA4 bit~0.4 bit~0.8 bit~0.4 bit~9.2GB

All three rows list activations as roughly 1–2× the parameter count, independent of method. The way to read the table: the expensive part of full fine-tuning isn't the weights. It's the gradients plus Adam's two moments, and those are exactly what LoRA cuts.

Page 7 sorts PEFT into three families, following the survey by Lialin et al.:

  • Selective: fine-tune a chosen subset of parameters.
  • Reparameterization: represent the weight update in low rank. LoRA lives here.
  • Additive: add new trainable layers or parameters, such as adapters or soft prompts.

LoRA: freeze W, learn A·B

The formula on pages 9–10 is one line. Freeze the pretrained weight W and learn a low-rank increment:

W′ = W + A·B, where A is d×r, B is r×d, and r is much smaller than d.

Page 10 also records where the idea came from. It first appeared in the multilingual translation paper Counter-Interference Adapter (CIAT, EMNLP 2021), and was later "re-invented" by LoRA (ICLR 2022). Their first arXiv versions are from April and June 2021. From here on the slides write "LoRA/CIAT".

Which matrices get it. Page 11 recommends rank 8 or 16, applied to attention's W_Q, W_K, W_V, and W_O, and not to the FFN's linear layers. The slide's reason is that the FFN is "storing knowledge". The same page adds that CIAT also applies it to embeddings and FFN layers, and that this helps. Read together, "attention only" is a default, not a law.

At inference (page 12). Compute W′ = W + A·B and use it as an ordinary weight; the only extra storage is A and B. To switch tasks, subtract the old adapter and add the new one: W″ = W′ − A·B + A″·B″. One base model can carry many small domain-specific adapters.

Backprop needs only two small gradients (page 13). W₀ is fixed, so you differentiate only with respect to A and B:

  • ∂L/∂A = g_out · (Bx)ᵀ
  • ∂L/∂B = (g_out · A) · xᵀ (as written on the slide; with the transposes spelled out by dimension it is (Aᵀ · g_out) · xᵀ)

If you did HW2 MiniTorch in this series, this is the matmul backward rule applied twice.

What you store (page 14): the original parameters, plus adapter weights, adapter gradients, and the adapter's two Adam moments (each 2×d×r), plus activations. The original parameters need no gradients or optimizer state.

Page 15 puts numbers side by side. LLaMA 8B's 8 billion parameters take 16GB in BF16, and Adam's optimizer state takes 48GB. With LoRA, trainable parameters drop to about 4M, which means 8MB of weights and 24MB of optimizer state. Page 16 concludes that memory and training time both go down, and inference works like any other LLM.

How large should the rank be

Pages 17–21 summarize the LoRA paper's experiments. NLU uses RoBERTa and DeBERTa on eight GLUE subtasks; NLG uses GPT-2 and GPT-3. Baselines are full fine-tuning, BitFit, two kinds of prefix tuning, and adapter tuning. The slides pull out two points:

  • Page 19: raising the rank doesn't cover more meaningful subspaces. A low-rank matrix with r=8 is enough.
  • Page 20: in a stress test scaled to 175B-parameter GPT-3, not every method improves monotonically as trainable parameters grow.

Page 21 also shows a chart of LoRA on GSM8K math problems. The slide gives no text stating a conclusion, so this post doesn't draw one for it.

QLoRA: put the frozen weights in 4 bits too

Page 23 separates two kinds of quantization. The previous lecture's GPTQ is post-training quantization: convert a trained model to lower precision, with no retraining. QLoRA is filed under quantization-aware training: quantization happens during training, which usually performs better.

Page 24 lists the three innovations of QLoRA (NeurIPS 2023): 4-bit NormalFloat storage, double quantization, and paged optimizers, with all computation in BF16. The slide cites the paper's result: average memory for fine-tuning a 65B model drops from 780GB to 48GB on a single GPU, with performance on par with 16-bit fine-tuning.

The algorithm on page 25 is a "store in 4 bits, compute in 16 bits" pipeline:

  1. Store the weights in NF4 with double quantization.
  2. Dequantize to BF16 when they're needed.
  3. Run forward and backward in BF16.
  4. Compute BF16 weight gradients only for the LoRA parameters.

NF4: why not equal-width bins

Four bits give you 16 values. Page 26 points out that absmax-style equal-width bins do badly on unevenly distributed data: lots of small values that differ only slightly get quantized into the same bin.

NF4 on page 27:

  1. Group weights in blocks of 64.
  2. Take the block's max absolute value as the scale.
  3. Divide the 64 numbers by the scale so they fall in [−1, 1].
  4. Find the nearest value in a 16-entry lookup table and output its index (−7 to 8).

The table isn't evenly spaced: bins are dense near 0 and sparse near ±1. Page 29 says the values come from quantiles ("probably quantile").

Pages 28 and 30 work one example by hand. It's worth following along:

  • Original value 0.0045, block scale 0.0071, normalized to about 0.63.
  • The nearest table entry is 0.5626, index 6.
  • Dequantized: 0.5626 × 0.0071 ≈ 0.00399, an error of about 0.0005.

Double quantization: the scales need compressing too

Page 31 names the cost of small blocks. One 32-bit constant per 64 parameters adds 32/64 = 0.5 bit per parameter on average.

Page 32 quantizes again. Each block of 64 values gets NF4; the resulting scales are grouped 256 at a time and quantized to FP8; only the second-level scale is stored in FP32. The slide's total is about 0.52 bytes per parameter.

Page 33 collects the full QLoRA configuration: pretrained weights in NF4 with double quantization, LoRA weights in BF16, and inputs plus all computation (forward, backward, optimizer) in BF16.

Paged optimizers: spill to CPU before you crash

Pages 34–36 deal with occasional OOMs during training. Optimizer state is allocated in NVIDIA unified memory. When the GPU runs short, those pages move to the CPU automatically, and they come back to the GPU when the optimizer update step needs them. The slides compare it to ordinary paging between CPU RAM and disk.

Pages 37–38 cite two results from the paper. LoRA's default hyperparameters don't match 16-bit performance, and NF4 beats 4-bit float (FP4). Once tuned, 4-bit QLoRA matches both 16-bit full fine-tuning and 16-bit LoRA.

Code walkthrough

Pages 40–42 are screenshots of a LoRA layer implementation, with three points: initialize the A and B layers; freeze the pretrained weight; in the forward pass, sum the outputs of the two branches (original weight and low-rank branch). After training, the LoRA weights can be merged back into the original weight.

Page 43 shows QLoRA usage: load the model with a quantization config, then apply LoRA as usual. The CUDA quantization ops come from wrappers in bitsandbytes. The slide's configuration looks like this (model name as on the slide):

AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-70b",
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=args.bits == 4,
        load_in_8bit=args.bits == 8,
        llm_int8_threshold=6.0,
        llm_int8_has_fp16_weight=False,
        bnb_4bit_compute_dtype=compute_dtype,
        bnb_4bit_use_double_quant=args.double_quant,
        bnb_4bit_quant_type=args.quant_type,
    ),
)

The same page links a Guanaco 7B Colab example in the QLoRA repo.

How this connects to the homework

LoRA shows up directly in HW6. The assignment has you turn on LoRA on top of DeepSpeed ZeRO so Llama-2-7B training fits on two 16GB V100s. The memory table from this lecture is the one you'll be doing in your head while tuning that run. For the ZeRO half, read L18 ZeRO first.

The summary on pages 44–45: LoRA/CIAT adapts large models cheaply with a small set of low-rank parameters, works at GPT-3 scale, and does better in low-data settings. QLoRA stores weights with NF4 double quantization and computes in BF16, matching the original LoRA. The slides call it the first method that can fine-tune a 33B LLM on a single consumer GPU.

How to study it yourself

  1. Read pages 4–6 and recompute the three rows yourself as "bits per parameter × parameter count". Make sure you know which part got cut.
  2. Follow pages 28 and 30 to quantize and dequantize with NF4 by hand, then repeat with a different number.
  3. Read the rank experiments in the LoRA paper and compare them with page 19.
  4. If you have a GPU, load a small model with the page 43 configuration and compare memory with bnb_4bit_use_double_quant on and off.

Further reading

References