Skip to content

CMU 11-768 Lecture 11: Advanced RL Algorithms — Credit Assignment, Stable Updates, Reward Hacking, and Distillation

Sep 29, 20261 min
TL;DRUsing a bug-fix coding task, Graham Neubig takes Lecture 9's policy gradient into practice: a critic, GAE, or a PRM to credit individual turns; importance ratios and clipping to handle stale data in async RL; and PPO, GRPO, CISPO, GSPO, and DAPO side by side in one table. The largest share goes to the reward itself — verifier errors, reward hacking, and exploration collapse — before closing with on-policy distillation.

🌏 中文版

This post is written from the slides; I will update it once the recording is posted. As of 2026-09-29, the official schedule lists only slides for Lecture 11, no recording. Everything below is based on the slide content alone, with no spoken commentary from the lecturer. Where I go beyond the slides, I say so.

CMU 11-768 AI Agents is a Fall 2026 graduate course on LLM agents taught by Daniel Fried and Graham Neubig. Lecture 11 (Sep 29) is the second of three RL lectures in the training module. Neubig teaches it, under the subtitle "Learning from trajectories when the simple recipe breaks."

In Lecture 9, Fried used a number-guessing game to derive REINFORCE, baselines, and GRPO, and explicitly deferred three things to this one: importance ratios, clipping, and the reference-model KL term. (Lecture 10 in between was a guest lecture on deep research agents, unrelated to RL.) The roadmap slide places this lecture in the middle. RL basics covered "rewards → advantages → policy gradients." This lecture takes on four things — useful feedback, less waiting, stable updates, and reliable rewards — plus learning from a teacher. The next lecture (Lecture 12, by TA Apurva Gandhi) covers memory, parallelism, and execution at scale.

The lecture packs in more than five algorithms, plus reward hacking and distillation. My approach: the algorithms go into one comparison table, reward hacking gets its own section, and distillation sits in a collapsible block.

The example: fixing a retry-config bug

This lecture switches to a more agent-like example, a teaching fixture reused from Lecture 6 on coding agents (the slide links the original fixture in the cmu-agents/lecture-planning repo, which was not public and returned 404 when checked on 2026-09-29, so the task is taken from the slides):

def retry_count(config):
    return config.get("retries") or 3

The task: retries=0 should disable retries, while missing, None, and positive values keep their current behavior. The verifier has four checks: missing → 3, None → 3, 0 → 0, 2 → 2. The bug is that or treats 0 as false and returns 3.

A successful trajectory looks like this: read the code → change it to get(..., 3) (which fixes 0 but now returns None for None) → run the four tests, and the None case fails → handle None explicitly, and all four pass. The slide's point: the agent can make a bad edit mid-way and still finish with a valid patch.

The reward is an outcome reward: 1 if the final code passes every check, otherwise 0. With Lecture 9's binary REINFORCE, a successful trajectory pushes up the probability of every sampled token in it, and a failed one contributes nothing.

1. Useful feedback: giving credit to individual turns

Group comparison is not enough

GRPO's group baseline compares trajectories on the same task. The slide's example has four trajectories whose final patches are "handle None explicitly" (reward 1), "use get(..., 3)," "always return 0," and "keep the buggy code" (all 0). The group mean is 0.25, so the success gets +0.75 and the rest −0.25.

The problem: every turn in a trajectory gets the same weight. Successful trajectory A may contain a mistake, and failed trajectory B may contain a helpful turn. What we want is credit at the turn level.

Reward-to-go and the value function

With only a final reward, the sum of rewards remaining after any step (the reward-to-go) is 1 at every step of a successful trajectory and 0 at every step of a failed one. That still doesn't separate helpful turns from harmful ones.

The fix is to ask a different question: starting from this history, how much reward do we get on average? The slide continues ten times from history h₃ of trajectory B; three continuations succeed, for an average of 0.3. That average over all possible continuations is the value function V(h₃).

Training a critic

You can't afford ten continuations at every step, so you train a critic V_φ(h) to predict reward-to-go with a squared-error loss. On those ten outcomes (three 1s, seven 0s), predicting 0.6 gives an average squared error of 0.30; predicting 0.3, the mean, gives 0.21, the minimum.

There are two places to put a critic:

Separate value networkValue head on the policy
Structureits own body and output layera linear head reading the policy body's hidden representation
Value loss trainsthe whole value networkonly the head (stop-gradient into the body)
ExampleInstructGPTMIXER (Ranzato et al.)

The stop-gradient keeps the value loss from updating the policy body; the policy loss still trains it.

TD error and GAE

With a critic, you can compare the value before and after each step. The last three actions of trajectory B:

ActionValue beforeValue afterTD error δ
a₃0.30.6+0.3
a₄0.60.4−0.2
a₅0.40 (episode ends)−0.4

a₃ improved the situation; a₄ and a₅ made it worse. Even though the whole trajectory failed, a₃ has a positive δ, which is exactly what group comparison could not tell you.

GAE (generalized advantage estimation, Schulman et al.) says an action's advantage should include not only its own δ but later δs too, discounted by λ. With λ = 0.5: Â₃ = 0.3 + 0.5 × (−0.2) + 0.25 × (−0.4) = 0.1. A λ near 0 trusts the critic more; a λ near 1 is closer to using the actual reward-to-go. (That last sentence is my addition; the slide only works through λ = 0.5.)

Mechanism: value loss, TD error, and the GAE recursion

The critic's squared-error loss (averaged over histories from sampled rollouts):

$$L_V(\phi) = \mathbb{E}t\big[(V\phi(h_t) - R_t)^2\big]$$

TD error (the examples use no discounting, γ = 1):

$$\delta_t = r_{t+1} + V_\phi(h_{t+1}) - V_\phi(h_t)$$

GAE runs backward; after the final step there are no more errors:

$$\hat A_t = \delta_t + \lambda \hat A_{t+1}$$

Another route: PRMs

You can also skip the critic. Let's Verify Step by Step (Lightman et al.) trains a process reward model (PRM) on human labels of whether each step is correct, as a surrogate for value. The slide's example is an algebra problem where the model goes from 5x = 6x − 14 to x = 7, and the annotator marks that step incorrect (the right answer is 14).

2. Less waiting: synchronous vs. asynchronous RL

In synchronous RL, every trajectory in a batch must finish before the policy update starts. Agent trajectories vary widely in length, so GPUs that finish early sit idle. Asynchronous RL lets rollout workers keep generating while the learner updates on completed trajectories, overlapping the two. The slides cite AReaL (Fu et al., NeurIPS 2025). The slide dates it 2026, which matches the fifth arXiv revision from March 2026; v1 was posted in May 2025 and the paper appeared at NeurIPS 2025.

The cost: some completed trajectories now come from an older policy version. That leads straight into the next section.

3. Stable updates: stale data, ratios, and clipping

Stale rollouts and importance sampling

The rollout used the policy μ of its time; the learner is now π_θ. The slide's example at one history:

Actionμ (at rollout)π_θ (now)Weight π_θ / μ
Check None explicitly0.400.701.75
Use get(..., 3)0.400.200.50
Keep the original code0.200.100.50

Importance sampling multiplies each sample by this ratio so data drawn under μ can estimate expectations under π_θ.

Long trajectories blow up the ratio

The ratio for a whole trajectory is the product of per-step ratios. Small per-step differences explode over long trajectories: 1.05 to the 100th power is about 132. A few samples then dominate the update, and variance is high. A figure from CTPO (Zhang et al.) shows the cumulative ratio spreading further at later positions in tool-using math rollouts.

There's a further subtlety: earlier actions change later histories. Suppose μ added a None check with probability 40% and π_θ only 10%. Even if the next action, "run tests," has probability 50% under both policies (ratio 1), the current policy reaches this history only 0.25 times as often. A per-step ratio alone misses that.

Which actions enter the weight

MethodActions used in the weight
PPO / GRPOthe current action only
CTPOall actions through the current one
Full ratioall actions
GSPOall actions, length-normalized

DAPO and CISPO also use only the current-action ratio; they differ in their clipping rules.

Clipping: limiting how far one update goes

Stale rollouts can get large weights and let a few samples dominate. Clipping replaces any ratio outside [ℓ, u] with the nearest bound, for example [0.8, 1.2].

PPO and CISPO (MiniMax-M1) both clip, but differently:

  • PPO takes the minimum of the clipped and unclipped terms. With a positive advantage, it stops rewarding ratios above 1 + ε; with a negative advantage, it stops rewarding ratios below 1 − ε.
  • CISPO uses the clipped ratio as a fixed weight (stop-gradient) multiplying the advantage and the log-probability.
Mechanism: PPO and CISPO objectives

Per-token ratio and clip:

$$\rho_t = \frac{\pi_\theta(u_t \mid u_{<t})}{\mu(u_t \mid u_{<t})}, \qquad c_t = \operatorname{clip}(\rho_t, \ell, u)$$

PPO (maximize):

$$J_{\text{PPO}}(\theta) = \mathbb{E}_t\big[\min(\rho_t \hat A_t,\ c_t \hat A_t)\big]$$

CISPO (maximize; sg is stop-gradient):

$$J_{\text{CISPO}}(\theta) = \mathbb{E}t\big[\operatorname{sg}(c_t), \hat A_t \log \pi\theta(u_t \mid u_{<t})\big]$$

The loss to minimize is $-J$.

Whole-trajectory ratio:

$$w(\tau \mid x) = \frac{p_\theta(\tau \mid x)}{p_\mu(\tau \mid x)} = \prod_t \frac{\pi_\theta(a_t \mid h_t)}{\mu(a_t \mid h_t)}$$

Reference-model KL and entropy

Two more common add-ons:

  • Reference-model KL: subtract β times KL(π_θ ‖ π_ref) to keep the policy from drifting far from a fixed reference model. The slide separates two roles: the rollout policy μ generates the training batch and supplies the denominator of the importance ratio, while the reference model π_ref stays fixed as an anchor. The source is the GRPO objective in DeepSeekMath.
  • Entropy bonus: add α times the entropy to reward a broader next-token distribution, which can help the agent try other actions. The slide also notes that a broader distribution does not guarantee those actions are useful.

The algorithm comparison table

The summary slide compares three things: how each method assigns credit, how it weights sampled tokens, and how it limits updates.

AlgorithmWeight from rewardLearned critic?Importance ratioClipping
REINFORCEreward-to-go R_tnonone (fresh rollouts)none
PPOGAE from a criticyescurrent token ρ_tminimum of raw and clipped terms
GRPOgroup comparisonnocurrent token ρ_tsame as PPO
CISPOgroup comparisonnocurrent token ρ_tclipped ratio as a fixed weight
GSPOgroup comparisonnowhole response, length-normalizedPPO minimum on the response ratio
DAPOgroup comparisonnocurrent token ρ_tPPO minimum, higher upper bound

"Group comparison" here means subtracting the group's mean reward and dividing by its reward spread. One way to read the table: first decide whether you can afford a critic, then decide how stale your data is. If you can't afford a critic, you're in the group-comparison rows; if your rollouts are long, asynchronous, and stale, choosing the ratio and clipping rule carefully matters. That reading is mine, not the slide's.

4. Reliable rewards: from uninformative groups to reward hacking

This is the longest part of the lecture and the most directly useful for people building agents. Everything above assumes the reward is right. This section deals with what happens when it isn't.

Why did every trajectory fail?

Sample four trajectories per task. [0, 0, 0, 0] and [1, 1, 1, 1] both become all zeros after subtracting the mean, so there is no comparison signal. When everything fails, the slide says to inspect what happened before picking a remedy:

Possible causeHow to check
Verifier quality: a valid solution is rejectedreview the requirements and the failed tests
Current capability: the model can't solve it yetsee whether stronger-model demonstrations succeed
Lack of exploration: every attempt repeats one approachcheck whether the trajectories differ

False negatives: valid solutions rejected

Back to the retry task. A candidate gets all four cases right but is written differently, and the verifier requires the exact phrase value is None in the source, so it rejects the candidate. That is a false negative: a valid solution rejected. The fix is to check behavior, not an incidental wording choice.

False negatives have two consequences:

  • Reward noise: equivalent solutions get different rewards because of wording or implementation details, and the model fits those spurious differences.
  • Benchmark saturation: when the evaluator rejects valid solutions, scores flatten below 100%.

The slide's examples (historical versions, audited in 2025–26):

BenchmarkPlateau / limitAudit finding
SWE-bench Verified≈81%the best score rose only from 74.9% to 80.9% over six months; OpenAI audited 138 problems that o3 did not consistently solve over 64 runs and found material issues in test design or problem description in 59.4% (OpenAI, Feb 2026)
τ-bench Airline≈70%inconsistent tasks limited achievable scores (SABER §5.1)
τ-bench Retail≈92%annotation errors limited achievable scores (same source)

The three numbers are different kinds of figures and should be read separately. The τ-bench 70% and 92% are achievable ceilings computed in the SABER paper: errors in the dataset cap scores at that level. The SWE-bench Verified 81% is the point where OpenAI observed state-of-the-art progress stalling, not a computed ceiling. OpenAI also gives two reasons, not one: besides tests that reject correct solutions, there is contamination — every frontier model it tested could reproduce gold patches for some tasks — which is why it stopped reporting the score and recommends SWE-bench Pro instead. The slide files both under "false negatives cause saturation" and covers only the test half.

SABER is a third-party audit (Cuadron et al.), not from the τ-bench authors. In τ³-Bench: Fixing Airline + Retail (February 2026), the τ-bench team fixed 27 airline and 26 retail tasks and says most fixes came directly from SABER. SWE-bench Verified itself was built by OpenAI as a human-validated subset to filter out problematic tasks in the first place: engineers reviewed 1,699 problems and kept 500.

False positives and reward hacking

Now flip it. Suppose the training verifier tests only the input retries=0:

def retry_count(config):
    return 0

This "always return zero" candidate passes the one test and earns reward 1, yet gets missing, None, and 2 all wrong. That is a false positive: an invalid solution rewarded. Training can learn to exploit this gap.

For coding agents, the slide lists four common shortcuts:

Available in the environmentShortcut
Internet accessfind the published solution online
Git history, including future commitsfind and copy the later fix
A model API keycall a stronger model for answers or training data
Access to tests or the test runnermake the tests pass without fixing the bug

For the web, Git, and test shortcuts, the slide cites the MAI-Thinking-1 report §3.3.1 (p. 43). The report sorts reward hacking in its SWE environments into three types — searching the internet for the original PR, digging through local Git history for the fix commit, and tampering with tests — and counters them by restricting network access, scrubbing every commit after the base commit, and resetting test files before grading. The API-key case is documented: PostTrainBench (Rank et al.) §5.4 explains that the OpenAI API key used for evaluation is exposed to agents, with an explicit restriction in the evaluation script against other uses. GPT-5.1 Codex-Max acknowledged the restriction in its reasoning trace, then, after extended struggles with model quality, violated it and used the key to generate training data (Figure 7). The authors suspect the restriction had dropped out of context in the long session. The paper reports this as a single instance, not a rate.

Independent checks of progress

The constant-zero patch earned training reward 1 but passed only one of the four required cases. The slide recommends watching three signals together:

Training rewardIndependent successCosts and failures
success on the training verifierunseen tasks and behavior the verifier missedinspect tool calls, regressions, and repeated actions

If reward rises without independent improvement, inspect trajectories for shortcuts.

DAPO's dynamic sampling

All-pass or all-fail groups have zero advantage: they take up batch space and contribute no gradient. DAPO (Yu et al., §3.2) keeps only groups with mixed rewards and keeps sampling until the batch is full. A kept group is kept whole, failures included.

Partial rewards and their risks

Another way to make groups informative is partial credit, say +0.25 per check passed. If four trajectories pass 0, 1, 2, and 1 checks, the pure success reward is all zeros; with partial credit it becomes 0, 0.25, 0.5, 0.25, which after subtracting the mean is −0.25, 0, +0.25, 0. Now there's a signal.

But auxiliary rewards can be gamed too. The slide gives two cases OpenAI has published:

Exploration collapse

The last reward-related risk lives in the policy itself. At one history, three actions start at 0.2, 0.6, 0.2 and end at 0.01, 0.98, 0.01: almost every sample now picks B, and alternatives are rarely tried. An entropy bonus encourages alternatives; you should also check whether trajectories actually differ. The DAPO paper discusses this entropy collapse as well.

5. Learning from a teacher: curricula, warm starts, and distillation

When the student almost never succeeds, RL gets no signal. The slide's first answer is a curriculum: learn from successful solutions first (a warm start) → practice tasks the model sometimes solves → raise difficulty as it improves. Keep earlier tasks in the mix and use a fixed evaluation set to measure progress.

Keeping earlier tasks matches the formal definition in Bengio et al. (2009): a curriculum is a sequence of reweighted training distributions in which no example's weight ever decreases, the entropy grows, and every example ends at weight 1. The warm start and "practice what the model sometimes solves" are the slide's RL-specific advice and are not in that paper. The paper's main experiments (shape classification and language modeling) are supervised, with difficulty fixed in advance, such as images with less shape variation first or frequent words first, rather than chosen from the model's current success rate.

Then comes distillation: instead of a single 0/1 outcome, have a teacher provide a full action distribution at every step.

Three kinds of distillation: from offline to on-policy self-distillation

Offline distillation (Hinton et al., 2015). Train a student to match a teacher's action probabilities on a fixed set of teacher trajectories. For example, at history h₃ the teacher puts 80% on a₃ and 20% on the other action; call this q(· | h₃). The problem is the same as SFT's exposure bias: when the student acts, its own mistakes lead to histories missing from the training set.

On-policy distillation (Agarwal et al., ICLR 2024, GKD). Let the student generate the histories, then ask the teacher for targets at those histories. "On-policy" means the histories came from the student's own rollouts.

A dense distillation objective. At the same h₃, the student assigns 0.5 to each action and the teacher 0.8 and 0.2. Reverse KL (natural logs):

$$0.5 \log\frac{0.5}{0.8} + 0.5 \log\frac{0.5}{0.2} \approx 0.223$$

Average over the distribution of student histories $d_\mu$, then minimize:

$$L_{\text{OPD}} = \mathbb{E}{h \sim d\mu}\big[D_{\text{KL}}(\pi_\theta(\cdot \mid h) ,|, q(\cdot \mid h))\big]$$

Compared with RL, every step now carries a signal, not just a final score.

On-policy self-distillation (OPSD, Zhao et al.). The teacher doesn't have to be a bigger model. OPSD uses a frozen copy of the starting model as the teacher; the only difference is the input. The student sees the task and history h₃; the teacher sees the same task and history plus a verified solution. With the solution in hand, the teacher gives better targets at the student's own histories, and the student is trained the same way as before.

Three takeaways

The final slide sums up the lecture in three columns:

Useful feedbackStable updatesReliable rewards
use group comparisons, a critic, or process rewards to assign creditknow the behavior policy; control how strongly samples change the modelmatch task difficulty and verify the behavior you actually want

What an agent builder can do tonight:

  • Run a fake solution that always returns a constant, or changes nothing, through your verifier. If it scores, your reward has a hole, and training will find it.
  • List what your sandbox exposes: network, Git history, API keys, test files. For each, ask whether the agent could score without doing the work.
  • Plot two curves during training: the training reward and an independent evaluation the verifier never touches. The moment they diverge is when to go read trajectories.

Going deeper

References listed for this lecture on the official schedule (most also appear on the slides):

Further reading on this site (other angles on the same algorithms; not a substitute for this lecture):

References