Skip to content

MIT 6.5940 L14 LLM Post-Training: From SFT and RLHF to Fine-Tuning That Touches 1% of the Weights

Sep 30, 20261 min
TL;DRLecture 14 has three parts. Fine-tuning: SFT runs next-token prediction on desired answers, RLHF trains a reward model and then fine-tunes with KL-penalized RL, and DPO collapses both stages into one supervised step. Then comes a chain of PEFT methods: BitFit tunes only biases, Adapters add small layers but slow inference, Prompt/Prefix-Tuning eat input length, LoRA fixes latency with a low-rank branch you can merge back, QLoRA stores the backbone in NF4, and BitDelta compresses the fine-tune delta to 1 bit. Multimodal LLMs: Flamingo uses cross-attention, PaLM-E and VILA feed images in as tokens, and VILA-U can also output images. Prompt engineering: zero/few-shot, CoT, and RAG.

🌏 中文版

This post is based on MIT 6.5940 Fall 2024. It is post 18 in the Reading MIT 6.5940 series.

Series: previous Fall 2026 supplement: Lab 1 GPU Basics | next L15 Long-Context LLM | Series overview

Official materials: Lec14-LLM-Post-training.pdf (94 pages; all page numbers below refer to this PDF) and the Lecture 14 recording. The F24 schedule puts this lecture on October 24, 2024, the same day the final project ideas came out. Access level A3: slides and video are public, and no lab goes with this lecture. Checked on 2026-09-30.

A heads-up: the PDF cover says "Lecture 13 LLM Post-Training Part II", and page 26 contains an older, different Lecture Plan (mentioning PockEngine and LongLoRA). The course page and the video title both call this Lecture 14. This post follows the course page and uses the Lecture Plan on page 3 as its map.

Fall 2026 comparison: The Fall 2026 course page also schedules "LLM Post Training" (Lecture 14, October 29). As of 2026-09-30 its slide and video links are still empty, so there is nothing to compare yet.

What this lecture is about

Lecture 13 asked how to run a trained LLM fast. This lecture steps back: a pretrained model can't act as an assistant or read images yet. How do you turn it into what you need at the lowest cost?

The Lecture Plan on page 3 has three parts:

PartSlidesContent
1. LLM fine-tuning5–42SFT, RLHF (plus DPO), PEFT: BitFit, TinyTL, Adapter, Prompt-Tuning, Prefix-Tuning, LoRA, QLoRA, BitDelta
2. Multimodal LLMs44–73Cross-attention (Flamingo), visual tokens (PaLM-E, VILA), image output (VILA-U)
3. Prompt engineering75–93In-context learning, chain-of-thought, RAG

The course is about efficiency, so the PEFT sequence gets the most detail here. Each method fixes a problem the previous one left behind. Read in order, they form one clear line of development.

Part 1: Fine-tuning

SFT, RLHF, and DPO: three ways to teach a model how to answer

  • SFT (page 5): The objective is still next-token prediction. The only change is that the data contains only the answers you want. The slide uses Llama-2's SFT data, split into helpfulness and safety examples. Without SFT the answer is dry and terse. With it, the model sounds like customer support.
  • RLHF (pages 7–9): Cites Ouyang et al. (InstructGPT). Static metrics such as BLEU and ROUGE can't capture creativity, truthfulness, or usefulness, so RLHF optimizes for human preference directly. There are two steps. First, train a reward model on pairwise comparisons. Then use RL to maximize reward, with a KL term that keeps the model close to a reference model so it doesn't overfit the reward model.
  • DPO (pages 10–11): Rafailov et al. replace the two-stage, multi-model RLHF pipeline with one supervised step. No reward model and no RL algorithm. Page 11 uses the question "Where is Shanghai?": y_win is "Shanghai is a city in China" and y_lose is "Shanghai does not exist". The reference model's log probabilities can be computed offline in advance.
The three objectives (pages 8, 9, 10)

Reward model:

$$ \max_{r_\theta}\ \mathbb{E}{(x,y{win},y_{lose})\sim\mathcal{D}}\left[\log\sigma\big(r_\theta(x,y_{win})-r_\theta(x,y_{lose})\big)\right] $$

RL fine-tuning:

$$ \max_{\pi_\theta}\ \mathbb{E}{x\sim\mathcal{D},,y\sim\pi\theta(y|x)}\left[r_\theta(x,y)\right]-\beta,\mathbb{D}{KL}\left[\pi\theta(y|x),|,\pi_{ref}(y|x)\right] $$

DPO:

$$ \max_{\pi_\theta}\ \mathbb{E}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_{win}|x)}{\pi_{ref}(y_{win}|x)}-\beta\log\frac{\pi_\theta(y_{lose}|x)}{\pi_{ref}(y_{lose}|x)}\right)\right] $$

Other courses cover these three methods in more depth (see Further reading). 6.5940 spends seven slides on them and puts the weight on what comes next: whatever the training objective, full fine-tuning updates billions of parameters, and storing dozens of fine-tuned copies costs even more.

PEFT: each method patches the previous one

Page 17 makes the case for PEFT with one comparison. Store a full 7B LLaMA for each of 1000 downstream tasks and you need 14 PB. Store a 14 MB adapter per task and you need 14 GB.

In slide order:

MethodPagesHow it worksWhat it leaves unsolved
BitFit13–15Update only biases. BERT-base has 110M parameters but only 0.1M biases, over 1000x fewerMatches or beats full fine-tuning on small-to-medium datasets; falls behind with more data
TinyTL16Tune only biases, and add lite residual modules for capacity. Keep activations small: lower resolution, avoid inverted bottlenecksBuilt for on-device training; Lecture 21 goes deeper
Adapter17–19Insert small bottleneck layers into each Transformer layer and train only those. New tasks don't touch old onesThe extra layers sit in series at inference time and add latency
Prompt-Tuning20–22Replace a hand-written prompt with a trainable continuous vector prepended to the input. One batch can mix prompts for different tasks. Accuracy approaches full fine-tuning as models growApplied only at the first layer
Prefix-Tuning23–24Add trainable prefixes at every layer; consistently beats embedding-only tuningSee below

Page 25 spells out what Prompt-Tuning and Prefix-Tuning share: both lengthen the input. That slows inference and uses up sequence length you could otherwise use. The slide then poses the question for the whole section: can we fine-tune without adding any inference latency?

LoRA: a branch during training, merged away at inference

Intuition. Adapters are slow because the extra layers sit in series on the main path. Put the branch in parallel instead, and make it something you can add back into the original weight matrix. Then inference looks as if nothing changed.

Mechanism (pages 27–30). LoRA adds two small matrices beside each layer. A projects dimension d down to rank r and is initialized from a Gaussian. B projects r back to d and is initialized to zeros. Because B starts at zero, adding the branch doesn't change the output at first:

$$ h = xW + xAB = x(W + AB) = xW' $$

After training, add $AB$ into $W$ to get $W'$. Inference costs nothing extra.

Pages 31–35 show LoRA outside language models. The same stable-diffusion-v1-5, with different LoRAs downloaded from Civitai, switches to ink-wash scenery, detail enhancement, and other styles.

QLoRA and BitDelta: plugging quantization into fine-tuning

If you've done Lectures 5–6 on quantization, these two are easy to follow:

  • QLoRA (pages 36–39): Keep LoRA's design, but store the frozen backbone in 4 bits. Three parts: a new data type, NormalFloat (NF4; page 37 lists its 16 exact values); double quantization, which also quantizes the scaling factors; and paged optimizers with CPU offloading. The slides conclude that mid-range and entry-level GPUs can now fine-tune LLMs.
  • BitDelta (pages 40–42): The intuition is that fine-tuning adds little new information, so the weight delta should compress well. BitDelta quantizes the delta to 1 bit and fine-tunes one scaling factor per tensor. Page 41 describes a fused binary GEMM kernel that combines dequantization with the matrix multiply, so 1-bit deltas stay quantized during batched inference. Page 42 applies it to multi-tenant serving, with all models fine-tuned from Mistral-7B. The slide's tagline: "The more you serve, the more you save!"

The whole PEFT line fits in one sentence: first shrink what you train (BitFit, Adapter, Prompt), then remove the inference cost (LoRA), then cut storage and serving cost too (QLoRA, BitDelta).

Part 2: Multimodal LLMs

Page 45 splits the ways to make an LLM see into two camps:

  1. Inject vision through cross-attention (the Flamingo approach)
  2. Feed visual tokens as input (the PaLM-E approach)

Flamingo: freeze the LLM, insert cross-attention

Flamingo (pages 46–50) keeps the LLM frozen and inserts cross-attention layers between its layers so text can attend to images. Two components:

  • Perceiver Resampler (page 47): compresses variable-size image features into a few fixed visual tokens. The slide's example: 27 visual tokens plus 5 learned queries give 32 keys and values, a 5×32 attention map, and 5 output tokens.
  • Gated cross-attention (page 48): a tanh gate controls how much visual information gets in. The gate starts at 0, so the LLM behaves exactly as before at first. It's the same idea as initializing LoRA's B to zero.

PaLM-E and VILA: turn images into tokens

  • PaLM-E (page 51) feeds images, robot states, and 3D representations into the LLM as tokens. Page 52 continues to RT-2, which outputs control signals directly.
  • VILA (pages 53–64) comes from Song Han's lab and trains in three stages: projector training, pretraining, and SFT. Pages 54–57 list four findings:
    • Freezing the LLM during pretraining gives decent zero-shot results but no in-context learning. You need to unfreeze the LLM to get it.
    • Interleaved image-text data helps; image-text pairs alone are not enough.
    • Mixing text-only instruction data back in during SFT recovers text-only performance and also improves vision-language accuracy.
    • Original image resolution matters more than the number of tokens.

Page 64 shows VILA beating LLaVA-1.5 with the same prompts and the same base LLM. Pages 65–66 add two ways to handle high resolution: InternVL 1.5 tiles the image into 448×448 pieces plus a thumbnail, and CogAgent attaches a lightweight high-resolution encoder through cross-attention.

VILA-U: output images too

VILA-U (pages 67–73) puts understanding and generation of video, images, and text into one autoregressive model. Two keys:

  • A unified vision tower: trained with both an image-text contrastive loss (for semantics) and a reconstruction loss (to keep appearance, which generation needs), using residual quantization to turn images into discrete tokens.
  • Token in, token out: every modality becomes tokens. Training can apply the LM loss to any token, and at inference decoders turn tokens back into text, images, or video.

Page 70 claims this is the first time a VLM with discrete visual tokens matches continuous-token models on understanding. The "tokenize the image, then generate autoregressively" idea returns in Lecture 16's HART, this time from an efficiency angle.

Part 3: Prompt engineering

The last part needs no training. It's about how you ask:

  • Zero-shot (pages 75–76): Language models used to be one model per task, with one BERT for translation and another for sentiment. Once models got big enough, emergent abilities let one foundation model handle many tasks through prompts.
  • Few-shot / in-context learning (pages 77–78): put a few demonstrations in the prompt. Page 78 gives two practical tips. Balance the number of examples per class for classification. Keep the demonstration format consistent.
  • Chain-of-thought (pages 79–80): have the model write out intermediate reasoning steps. In zero-shot settings, adding "Let's think step by step" also works.
  • Prompts for diffusion (pages 84–91): SDXL examples show the effect of adding descriptors one at a time, plus a negative prompt. Page 83 also includes a "grandma jailbreak" example, marked as already fixed.
  • RAG (pages 92–93): an LLM can't memorize all long-tail knowledge. A simple RAG pipeline has four parts: an embedding model (evaluated with benchmarks like MTEB), a retriever, an optional reranker, and a language model that writes the answer.

How to self-study this lecture

  1. Make sure you understand the question on page 25 ("can we fine-tune without extra inference latency?") before reading $W' = W + AB$ on page 30. Together these two pages explain why LoRA beats Adapters.
  2. Redo page 17's 14 PB vs 14 GB comparison for your own situation. How many fine-tuned variants do you keep? What does each cost stored in full, and what does it cost as a LoRA?
  3. One thing you can do tonight: pick a LoRA checkpoint you've used (language or diffusion), open its config to find the rank r, estimate its parameter count as d×r×2, and compare that with the original layer.

Further reading

References