Skip to content

CS234 Learning from Human Preferences: Bradley-Terry, RLHF, and the DPO Derivation

Sep 30, 20261 min
TL;DRCS234 L8 keeps the inverse RL problem from the previous post but changes the input. Instead of expert demonstrations, a human says "A is better than B." The Bradley-Terry model turns these pairwise comparisons into a reward you can fit with cross-entropy. RLHF runs PPO on that reward model with a KL penalty. DPO shows that the KL-constrained optimal policy has a closed form and rewrites the reward as a log-ratio of policies. Plugged back into Bradley-Terry, the partition function cancels, so you can train the policy on preference data directly, with no reward model. The 2026 slides contain no offline RL, and DPO is now taught in lecture rather than by 2024's guest speakers.

🌏 中文版

Edition note: this guide follows the CS234 Winter 2026 Lecture 8 slides pp.18–68 and Lecture 9 pp.2–3 (PDF page numbers). L8 pp.37–58 are credited as slides from Tatsu Hashimoto's Lecture 11 in CS224N. The 2026 recordings are for enrolled students only. Videos 8 and 9 of the public Spring 2024 playlist serve only as supplements. Access level A3 (defined in the global AI/CS course map). Every fact was checked on 2026-09-30 against those PDFs and video pages.

Series: previous Learning from demonstrations: BC, DAgger, IRL, MaxEnt IRL | next A3: Hopper reward engineering, preference learning, DPO, best arm identification | Series overview

The previous post ended with a line from the L7 summary: often we only have preference pairs, not demonstrations. L8 starts there. Its plan reads "Imitation Learning and RLHF and maybe DPO," and the next class meeting is the midterm.

How 2026 differs from 2024

Two things first, so you do not hunt for missing material in the 2024 videos:

  1. The 2026 slides contain no offline RL. "Offline RL" appears in the labels for Weeks 4–6 on the 2026 schedule, but no lecture in the L6–L9 slides covers it. Playlist video 8 is titled "Offline RL 1," but its YouTube chapters show MaxEnt IRL and RLHF, so the table below uses it as a supplement. The lecture that actually covers offline RL is video 10, "Offline RL 3," which has no matching 2026 slides, and this post does not invent offline RL content for 2026.
  2. DPO moved from a guest lecture into the main lecture. 2024 video 9 is a guest lecture by the DPO authors Rafael Rafailov, Archit Sharma, and Eric Mitchell. In 2026, DPO is part of L8 and uses the CS224N slides.
2026 L8 slidesContentSpring 2024 supplement (per YouTube chapters)
pp.18–25How humans can help train RL agentsVideo 08 from 52:26, "RLHF techniques"
pp.26–35Pairwise comparisons, Bradley-Terry, trajectory rewards, the backflipVideo 08 from 1:00:36, "Bradley-Terry model"
pp.36–44RLHF pipeline, reward models, InstructGPTVideo 08 1:11:10, "RLHF in LLMs"; video 09 from 5:02, "RLHF overview"
pp.45–64The DPO derivation and resultsVideo 09 from 26:21, "DPO theory introduction" to 52:19, followed by Q&A

How humans can help

L8 pp.20–25 widen the frame. If we want RL agents that match human performance or human values, human input matters. The slides give several examples:

  • Thomaz & Breazeal (2008) studied how people teach robots, to build better robot learners
  • A spectrum ordered by human effort: DAgger-style constant teaching at one end, demonstrations only at the other, and pairwise labels in between. The slides ask whether that middle is the "sweet spot"
  • Comparing recommendation rankings (from Yisong Yue's dueling bandits lecture)
  • Sadigh et al. (RSS 2017), who actively choose which comparisons to ask people, to learn rewards for human-robot interaction

Why pairwise comparisons? L8 p.27 gives two reasons. For people, comparing is usually easier than writing a reward function by hand. It is also easier than giving a scalar score ("how much do you like this ad?").

Bradley-Terry: turning "A beats B" into a probability

The slides first recall the previous lecture's conclusion: without further assumptions, the latent reward model is not unique. So they commit to one structural model.

Start with the simplest setting, a k-armed bandit with K actions b₁…b_K and no state. Assume a person makes noisy pairwise comparisons and prefers bᵢ over bⱼ with probability:

P(bᵢ ≻ bⱼ) = exp(r(bᵢ)) / (exp(r(bᵢ)) + exp(r(bⱼ)))

This is the Bradley-Terry model (1952). It is transitive: p_ik is determined by p_ij and p_jk.

Three ways to define "best"

L8 p.29 lists three winners from the dueling-bandit literature:

NameDefinition
Condorcet winnerBeats every other option with probability above 0.5
Copeland winnerMost wins minus losses
Borda winnerHighest expected score against all options (win 1, tie 0.5, loss 0)

The slides add that k-armed and dueling bandit algorithms have historically aimed at the Copeland winner. The table is a reminder that once preferences need not be transitive, "the best answer" itself needs a definition.

Fitting the parameters

The data are N tuples (bᵢ, bⱼ, μ): μ = 1 if the person marked bᵢ ≻ bⱼ, 0.5 for a tie, and 0 otherwise. Maximize the likelihood with cross-entropy:

loss = −Σ [μ log P(bᵢ ≻ bⱼ) + (1 − μ) log P(bⱼ ≻ bᵢ)]

This is essentially logistic regression, except that the input is the difference between two rewards.

From actions to trajectories

The same model applies to trajectories. Let R¹ be the latent, unobserved sum of rewards along trajectory τ¹. The probability that a person prefers τ¹ ≻ τ² is exp(R¹) / (exp(R¹) + exp(R²)). Once the reward model is learned, the slides' next step is a single line: use the learned reward model and run PPO.

That is the recipe of Christiano et al. 2017. The result quoted on L8 p.35 is a simulated robot learning to backflip, which "needed 900 bits of feedback from a human evaluator."

From backflips to ChatGPT: the RLHF pipeline

L8 pp.37–44 switch to CS224N slides and move the same pipeline to language models:

  1. Step one is instruction tuning
  2. Steps two and three maximize reward; the question is how

Why ask for comparisons: human scores are noisy, and different people's scales are miscalibrated. Asking "which is better" is usually more reliable. The reward-model loss is Bradley-Terry: push the "winning" sample's score above the "losing" one and take −log σ(difference).

Check the reward model first: citing Stiennon et al. 2020, the slides evaluate the reward model on held-out human judgments. A large enough reward model trained on enough data approaches single-human accuracy.

Putting it together: you have a pretrained (possibly instruction-tuned) model p^PT, a reward model, and a method that optimizes toward any reward. Copy the model into p^RL_θ and optimize with RL:

r(s) = RM(s) − β log(p^RL_θ(s) / p^PT(s))

The second term is a penalty. You pay a price whenever p^RL prefers an output more than p^PT does, and in expectation that penalty is the KL divergence between the two. It keeps the model from drifting too far from pretraining.

Three results follow:

  • Stiennon et al. 2020: RLHF beats pretraining plus fine-tuning alone
  • InstructGPT (Ouyang et al. 2022): scaling RLHF to roughly 30,000 tasks
  • Controlled comparisons (Dubois et al. 2023): many studies use GPT-4 feedback as a surrogate for human feedback. PPO, the method in InstructGPT, does work, but simple baselines such as best-of-n and training on "good" outputs also work well

L8 p.48 also cites Zheng et al. 2023, "Secrets of RLHF in Large Language Models Part I: PPO," on how troublesome PPO is to run on LLMs in practice.

DPO: you may not need the reward model

DPO (Rafailov et al. 2023) starts from the same RLHF objective: maximize reward under a KL constraint. L8 pp.49–58 reach a loss in four steps.

Step 1: write the objective. For any reward function r, maximize E[r(x, y)] − β KL(π ‖ π_ref).

Step 2: the optimal policy has a closed form. This result comes from prior work:

π*(y|x) = (1/Z(x)) · π_ref(y|x) · exp(r(x, y)/β)

Z(x) sums over all possible responses and is intractable. The slides note that this means you cannot use the formula directly.

Step 3: invert it. Take logs and rearrange to write the reward as a function of the optimal policy:

r(x, y) = β log(π*(y|x) / π_ref(y|x)) + β log Z(x)

The slides' reading: the ratio is positive when the policy likes a response more than the reference model does, and negative when it likes it less.

Step 4: plug into Bradley-Terry. Bradley-Terry only cares about the difference between two responses' rewards, so β log Z(x) cancels. What remains is a loss that contains only policies:

L_DPO = −E[log σ(β log(π_θ(y_w|x)/π_ref(y_w|x)) − β log(π_θ(y_l|x)/π_ref(y_l|x)))]

The slides sum it up as an equation: "a loss function on reward functions" plus "a transformation between reward functions and policies" equals "a loss function on policies."

Where the Step 2 closed form comes from (the direction of the handwritten work on L8 pp.54–55)

The handwritten page L8 p.54 starts from the closed form π*(y|x) = (1/Z(x)) π_ref(y|x) exp(r(x,y)/β) and solves for r. Divide both sides by π_ref and take logs to get log(π*/π_ref) = r/β − log Z, that is, β log(π*/π_ref) + β log Z = r. The slide continues: "now plug into Bradley Terry pref func, so we can derive pref directly in terms of π_ref and π*."

One way to see the closed form itself: expand the KL term, and the objective becomes minimizing the KL between π and the normalized distribution π_ref · exp(r/β). A KL divergence is smallest when the two distributions are equal. This paragraph is this guide's own explanation; the slides cite the result from prior work.

Results

L8 pp.59–64 show several results:

  • Reward vs. KL trade-off: generate positive IMDB reviews with GPT2-XL, use a pretrained sentiment classifier as the "gold" reward model, create preferences from it, and optimize with PPO and DPO. The comparison is how much reward each gets at a given KL.
  • Models trained with DPO: p.61 is a screenshot of a model leaderboard, with several DPO-trained models circled by hand
  • Large-scale DPO training: pp.62–64 excerpt the instruction fine-tuning sections of the Mistral and Llama 3 papers

p.65 is titled CPL, but in the post-version PDF that page has only the title and no content. CPL most likely refers to Contrastive Preference Learning (Hejna et al. 2023), another method that learns policies directly from preferences without RL. The slides do not expand on it, so this post only points to it.

The last two slides point further. Learning and deciding from human preferences sits where social choice, computational economics, and AI meet, and Stanford has a new course on it, Koyejo's CS329H "Machine Learning from Human Preferences." They also cite Knox & Stone's TAMER framework (2008), in which a human gives feedback while the agent trains.

The quiz that opens L9

L9, the bandits lecture, opens by testing this one. The question is "select all that are true":

  1. RLHF and DPO both learn an explicit reward model from preference data
  2. Both are constrained to be at most as good as the best examples in the pairwise preference data
  3. DPO does not use a reference policy
  4. None of the above
  5. Not sure

The post-version PDF does not mark the answer. After reading this post you can check it yourself. Compare statements 1 and 3 with Steps 3 and 4 of the DPO derivation. Statement 2 deserves more thought: RLHF's PPO stage generates new samples for the reward model to score. How does that differ from DPO, which trains only on fixed pairs?

How to self-study it

  1. Write the Bradley-Terry cross-entropy loss next to logistic regression and confirm that they differ only in taking a reward difference as input.
  2. While reading L8 pp.49–58, derive the DPO loss from the closed form yourself, and watch for the step where Z(x) disappears. If you get stuck, watch the "DPO loss function" chapter of 2024 video 09 (from 30:46).
  3. Compare with MaxEnt IRL from the previous post: P(τ) ∝ exp(wᵀμ_τ) and π* ∝ π_ref · exp(r/β) look alike. What plays the role of the reference distribution in each?
  4. Get hands-on in the next post's Assignment 3. Q2 learns a reward model from Hopper preference data and then runs PPO (run_rlhf.py). Q3 runs SFT and DPO on the same data (run_dpo.py).

One thing to do tonight: write a 5-armed Bradley-Terry simulator in numpy. Pick true rewards, generate a few hundred pairwise comparisons, and fit the rewards back with gradient descent. You will find the fitted rewards are correct only up to an additive constant, which is exactly why log Z can cancel in DPO.

Further reading

References