🌏 中文版
CMU 11-768 AI Agents is a Fall 2026 graduate course taught by Daniel Fried and Graham Neubig on LLM-based agents: tool use, planning, memory, training, safety, and human-agent interaction. Lecture 9 (Sep 22), taught by Fried, is the first of three RL lectures in the training module. This one covers the basic policy-gradient methods; Lecture 11 the following week, by Neubig, covers advanced algorithms and what it takes to make RL stable; Lecture 12, by TA Apurva Gandhi, covers RL systems and practical frameworks. (Lecture 10 in between is a guest lecture on deep research agents by Akari Asai.) Together they set up Assignment 3. The course site describes Assignment 3 only as implementing "the training procedures used to adapt and improve the agent"; Fried added in class that it will have you train models with RL in two environments: the Number Search game used throughout this lecture, and a simple crafting environment somewhat like Minecraft, needing more compute than the earlier assignments. Treat the environment details as provisional until the assignment is released.
This post is based on the recording and the slides. Fried opens by describing the lecture as a bridge from last week's SFT to policy gradients, which he calls the simplest and also one of the most elegant forms of reinforcement learning. The whole lecture compresses to one sentence: every method computes log-probabilities of the actions the agent itself took; they differ only in the weight they multiply by.
I follow the lecture's order and keep all the math in collapsible sections. If you only want the intuition, you can skip every "Mechanism" block and still follow along.
The setup: guess a number between 1 and 16
One example runs through the entire lecture. The environment hides an integer from 1 to 16 and the agent gets four guesses. After a wrong guess the environment replies "higher" or "lower." A correct guess ends the episode with reward 1; running out of guesses gives reward 0.
With the hidden number set to 11, the slides show four trajectories:
| Trajectory | Sequence | Outcome |
|---|---|---|
| τ₁ | 8 → higher → 12 → lower → 10 → higher → 11 | correct, R=1 |
| τ₂ | 4 → higher → 8 → higher → 12 → lower → 11 | correct, R=1 |
| τ₃ | 16 → lower → 8 → higher → 9 → higher → 10 | timeout, R=0 |
| τ₄ | 1 → higher → 2 → higher → 3 → higher → 4 | timeout, R=0 |
The best opening is 8 because it halves the range; Fried compares it to picking the optimal first word in Wordle. But τ₂ opened with 4 and still succeeded. τ₁ and τ₂ take different actions and reach the same outcome.
In agent terms, the reward could be whether the code passed every unit test, as on SWE-bench, or a non-binary score such as the fraction of tests passed. Much of today's RLVR work (reinforcement learning with verifiable rewards) uses 0/1 rewards, but everything in this lecture applies to any scalar reward.
The intuition: SFT has three gaps
SFT takes a demonstration and maximizes the probability the agent assigns to each demonstrated action. The demonstration can come from a person, from a stronger model, or from one of the agent's own successful trajectories. That last source is where this lecture starts.
Fried lists three problems with SFT:
Task mismatch. SFT maximizes the probability of the demonstrated actions. What we actually want is to maximize the probability of completing the task, and the two are not the same. First, there are many ways to succeed, and a demonstration shows only one. Second, the actions that were not demonstrated are not equally bad. Guessing 13 after "8 → higher, 12 → lower" is an awful move, yet SFT's loss penalizes it exactly as much as any other undemonstrated action.
Data mismatch. Even if the demonstration is optimal, sub-optimal or failed trajectories still carry signal. You would never run SFT on a timeout like τ₃, but you would like the model to learn not to do that. SFT has no way to express this.
Exposure bias. During SFT the history always comes from the demonstration. At inference time the history is whatever the model generated itself. If every training example follows the feedback correctly, the model has never seen the situation right after its own mistake. When it guesses in the wrong direction at inference time (say, 2 after "8 → higher"), it may have no idea how to continue.
What RL changes: the agent generates its own training data
In RL the agent generates its own trajectories, you score each one, and you adjust the parameters so high-reward trajectories become more likely and low-reward ones less likely. That closes all three gaps:
- The objective is expected reward itself, so any successful trajectory gets reinforced, not just the demonstrated one.
- Failed trajectories also give a signal to steer away from (once there is a baseline, covered below).
- Trajectories come from the current policy (on-policy), so the model faces its own mistakes in training and is penalized when they lead to low reward.
Fried adds an observation of his own (spoken in class, not on the slides, and given without a source): many of the repetition loops that language models got stuck in three or four years ago went away once people added an RL stage to post-training, which he attributes to models learning to recover from their own errors. This is the speaker's judgment; I did not find a study that directly tests this causal claim. The cost is efficiency. On-policy training has to keep waiting for the model to generate fresh data, and Neubig's Lecture 11 covers methods that relax this assumption.
Writing interaction as trajectories
The agent is a policy π_θ, where θ are the parameters of the underlying language model. At each step it sees the history h_t (every observation and action so far), samples an action a_t, and the environment returns the next observation and a reward. The slides map three settings onto this framework:
| Task instance x | Observation | Action | Reward | |
|---|---|---|---|---|
| Text generation | the prompt | tokens so far | next token | score from a judge model |
| Web agent | task + website | rendered page | click, type, scroll | number of sub-tasks completed |
| Number Search | the hidden number | higher / lower / correct | the next guess | 1 if correct, else 0 |
One distinction matters: the policy sees only the history, never the environment's hidden state. The hidden number 11 and the contents of a website's backend database both belong to the hidden state. When a student asked about the Markov assumption from classical RL, Fried explained that this is really a POMDP (partially observable Markov decision process): because the state is not visible, the next observation has to be conditioned on the full history.
A trajectory's probability is the product of two parts: the policy's probability of each action, and the environment's probability of returning each observation and reward. The key point in RL is that we never need to know the environment's part. We only need to interact with it and sample. That feels like magic, and it is also why RL is expensive: you can only try something and see what happens, then use that indirect signal to work out how reward depends on the parameters.
The objective is expected reward J(θ): sample trajectories under the policy, score them, and average. If the agent could only produce the four trajectories above, with probabilities 0.30, 0.20, 0.25, and 0.25, then J = 0.3 × 1 + 0.2 × 1 = 0.5. Training moves probability mass from the failing trajectories to the successful ones.
Mechanism: trajectories, trajectory probability, and the objective
History and policy:
$$h_t = (o_0, a_0, o_1, \ldots, o_t), \qquad \pi_\theta(a_t \mid h_t)$$
A trajectory with rewards, and its total reward:
$$\tau = (o_0, a_0, r_1, o_1, \ldots, a_{T-1}, r_T, o_T), \qquad R(\tau) = \sum_{t=1}^{T} r_t$$
By the chain rule, trajectory probability splits into a policy term and an environment term (in class Fried added the $P$ missing from the slide):
$$p_\theta(\tau \mid x) = \prod_t \pi_\theta(a_t \mid h_t), P(o_{t+1}, r_{t+1} \mid h_t, a_t, x)$$
Objective:
$$J(\theta) = \mathbb{E}{\tau \sim p\theta(\cdot \mid x)}[R(\tau)] = \sum_\tau p_\theta(\tau \mid x) R(\tau)$$
The simplest RL: SFT on the successful trajectories
Before getting to policy gradients, Fried introduces a simple, well-motivated modification of SFT: expert iteration, one form of which is ReST (Reinforced Self-Training, Gulcehre et al.). It repeats three steps:
- Grow: sample trajectories from the current policy and compute rewards.
- Improve: filter by reward (with binary rewards, keep R=1 and drop R=0), add the kept trajectories to the existing dataset, and run SFT on the whole dataset.
- Repeat with the updated policy.
In Number Search, τ₁ and τ₂ are kept and τ₃ and τ₄ dropped, so the loss is just the SFT loss on the successful trajectories. The EM in ReST-EM (Singh et al.) stands for expectation maximization, because this sample-then-refit alternation closely resembles the EM algorithm.
Fried stresses that this is not technically RL, but the skeleton is already the same: the model generates its own data, and the reward decides how that data updates it. It reuses the SFT machinery unchanged and is a simple, effective baseline when rewards are binary.
Two questions from class are worth keeping:
- How is this different from self-distillation? Only in using the reward to filter. Self-distillation can mean training on everything the model generates. ReST is "reinforced" because the reward is used to reinforce the good trajectories.
- What if the model never produces a successful trajectory? Then neither RL nor this method is a good fit. Fried gave three remedies: train on demonstrations first (the "cold start" that DeepSeek-R1 and many reasoning models do before RL); use a curriculum that starts with problems the model sometimes solves and adds harder ones as it improves; or add intermediate rewards, for example rewarding the model for exploring the repository first if you know that helps.
Policy gradient: the SFT loss times the reward
You cannot maximize J(θ) with plain backpropagation, for two reasons. Actions are sampled from a discrete distribution, so a small change in θ usually leaves the sampled action unchanged, and the derivative is zero almost everywhere. And the environment is a black box. As Fried puts it, how do you take the derivative of a unit test with respect to code?
The general answer is REINFORCE, proposed by Williams in 1992. An algebraic trick, the score-function (log-derivative) trick (Williams calls the ∂ln g/∂w term the characteristic eligibility), rewrites the gradient of expected reward as another expectation that can be estimated by sampling:
Sample a trajectory, sum the gradients of the log-probabilities of every action in it, and multiply by the trajectory's reward.
That final form carries two pieces of good news. First, the environment term is constant in θ, so it drops out of the gradient; you never need to know how the environment works. Second, what remains, the gradient of action log-probabilities, is exactly what SFT computes. The only difference is that these tokens were sampled by the policy instead of taken from a demonstration. When an action spans many tokens, such as a complex tool call, you just add up those tokens' log-probabilities.
Mechanism: deriving REINFORCE
Take the gradient of J. Only the trajectory probability depends on θ:
$$\nabla_\theta J = \sum_\tau \nabla_\theta p_\theta(\tau \mid x), R(\tau)$$
Multiply and divide by $p_\theta(\tau \mid x)$, and use $\nabla p / p = \nabla \log p$:
$$\nabla_\theta J = \sum_\tau p_\theta(\tau \mid x), R(\tau), \nabla_\theta \log p_\theta(\tau \mid x) = \mathbb{E}{\tau \sim p\theta}\big[R(\tau), \nabla_\theta \log p_\theta(\tau \mid x)\big]$$
The environment terms in the trajectory log-probability are constant in θ:
$$\log p_\theta(\tau \mid x) = \underbrace{\sum_t \log P(o_{t+1}, r_{t+1} \mid h_t, a_t, x)}{\text{const}} + \sum_t \log \pi\theta(a_t \mid h_t)$$
So one sampled trajectory gives a gradient estimate:
$$\hat\nabla_\theta J = R(\tau) \sum_t \nabla_\theta \log \pi_\theta(a_t \mid h_t)$$
When an action $a_t = (u_{t,1}, \ldots, u_{t,m_t})$ spans several tokens:
$$\nabla_\theta \log \pi_\theta(a_t \mid h_t) = \sum_{k=1}^{m_t} \nabla_\theta \log \pi_\theta(u_{t,k} \mid h_t, u_{t,<k})$$
Implementation: a mask and a scale
The slides lay out a Number Search trajectory as chat-template tokens: <|im_start|>assistant, guess(8), <|im_end|>, <|im_start|>tool, higher, and so on. Only the tokens the policy sampled (the guesses and the end-of-turn token) get mask m=1; prompt tokens and environment observations get m=0. The loss is:
$$L_{PG} = -R(\tau) \sum_{t : m_t = 1} \log \pi_\theta(u_t \mid u_{<t})$$
In other words, a policy gradient is the SFT loss, masked and rescaled. Fried shared a personal anecdote: he was surprised the first time he read REINFORCE code, from the Deal or No Deal paper (Lewis et al., 2017) that used RL to optimize negotiation dialogue agents. He expected something complicated; it was just the SFT loss multiplied by the reward. The anecdote is the speaker's own; the checkable part is that the paper's RL stage does use Williams's 1992 REINFORCE.
The training loop has four steps: sample task instances and let the agent generate trajectories → compute each trajectory's reward → weight action log-probabilities by reward → backpropagate, update θ, and sample again from the new policy. Standard REINFORCE draws one trajectory per task instance, so the batch size is the number of task instances.
With binary rewards, this looks a lot like ReST
When rewards are only 0 or 1, the REINFORCE loss equals the SFT loss on successful trajectories, the same as ReST. The difference is that REINFORCE is fully online and updates right after sampling, while ReST accumulates a large dataset and then runs one or more epochs over it. The intuition is the same.
Baselines and advantages: compare against usual performance
That is also where the trouble shows. τ₃ failed with reward 0, so its gradient is 0. Plain REINFORCE with no baseline, like ReST, learns nothing from failure, even though we clearly want the model to learn not to do that again.
Fried's intuition: the size of the update should depend on how well the policy usually does on this example. If the model scores 1 on a task 95% of the time, the 5% where it makes a silly mistake and scores 0 deserve a strong correction.
That is the advantage: how much better this action is than the policy's average. Formally it is the expected reward after taking this action minus the expected reward from this history, Q minus V.
A Number Search example: you have guessed 8 (higher) and 12 (lower), so the answer is 9, 10, or 11, with two guesses left.
- Guess 10: if it is wrong, the feedback tells you whether the answer is 9 or 11, and your last guess wins. Advantage > 0.
- Guess 11: if it is wrong, 9 and 10 remain with one guess, so you win only half the time. Advantage < 0.
In practice neither Q nor V is known, so we estimate the advantage as the sampled reward minus a baseline: Â_t = R(τ) − b(h_t). Your choice of baseline determines which algorithm you get:
| Baseline | Corresponds to |
|---|---|
| a constant | the bandit example below |
| a running mean of rewards | adapts as the policy improves |
| a trained value model V_φ(h) | actor-critic, PPO (Lecture 11) |
| the mean over several rollouts of the same task | GRPO, no extra model |
The slides present the baseline as an improvement on REINFORCE, but Williams's 1992 definition already includes it. The paper's update is Δw = α(r − b)e, where b is the reinforcement baseline, required only to be conditionally independent of the unit's current output. The name REINFORCE is an acronym for "REward Increment = Nonnegative Factor × Offset Reinforcement × Characteristic Eligibility," and "Offset Reinforcement" is r − b. The paper also gives an exponential moving average of past rewards as a baseline (reinforcement comparison, following Sutton 1984), which is the second row of the table above. "Plain REINFORCE" in this section means the b = 0 special case.
Why subtracting a baseline is safe
Subtracting a baseline changes the gradient computed from each sample, but not its expected value. As long as the baseline depends only on the history and not on the action chosen, the term it contributes has expectation exactly zero, so the estimate stays unbiased.
Mechanism: proof that baselines are unbiased
The estimator with a baseline:
$$\hat\nabla_\theta J = \sum_t \big(R(\tau) - b(h_t)\big), \nabla_\theta \log \pi_\theta(a_t \mid h_t)$$
It is unbiased as long as this term has expectation zero:
$$\mathbb{E}{a \sim \pi\theta(\cdot \mid h)}\big[b(h), \nabla_\theta \log \pi_\theta(a \mid h)\big] = b(h) \sum_a \pi_\theta(a \mid h), \nabla_\theta \log \pi_\theta(a \mid h) = b(h) \sum_a \nabla_\theta \pi_\theta(a \mid h) = b(h), \nabla_\theta 1 = 0$$
The second equality uses $\pi \nabla \log \pi = \nabla \pi$.
Worked example: a two-armed bandit
Fried says baselines always felt a bit mysterious to him, so he built a set of visualizations to see what they actually do. The setup is a single step with two actions: guess A scores 1, guess B scores 0. The policy has one parameter θ, and the probability of guessing A is p = σ(θ). At θ = −1.1, p = 0.25, so this is a bad policy that puts 75% of its probability on the zero-reward action. We want the gradient to push θ upward.
| sample A (prob. 0.25) | sample B (prob. 0.75) | |
|---|---|---|
| slope of log-prob w.r.t. θ | 1 − p = +0.75 | −p = −0.25 |
| no baseline: R × slope | 1 × 0.75 = +0.75 | 0 × (−0.25) = 0 |
| baseline b = 0.5: (R − b) × slope | 0.5 × 0.75 = +0.375 | (−0.5) × (−0.25) = +0.125 |
Both have expected gradient 0.25 × 0.75 = 0.1875, the true gradient. They differ in spread. Without a baseline, three draws in four give 0 and one gives 0.75, far above the true value; the standard deviation is about 0.325. With b = 0.5, both samples land near the true value and the standard deviation falls to about 0.108.
The sample-B cell is the most telling. Without a baseline you learn nothing. With one, "scoring 0 was worse than expected" is itself a signal: lowering B's probability raises A's. A baseline turns failures into a learning signal without changing the expected gradient.
A large part of RL work is reducing the variance of gradient estimates so training stays stable and isn't pushed around by sampling noise from the policy and the environment. The two-armed intuition carries straight over to multi-step decision problems.
What baselines do not solve
A baseline can mark a whole trajectory as better or worse than expected, and a good one cuts variance. It does not solve credit assignment: every action in a trajectory gets the same Â, so τ₁'s 8, 12, 10, and 11 all get +0.5. Estimating V may also require training a separate value model.
If the environment pays rewards mid-trajectory (say +0.2 for some tool call), replace R(τ) with reward-to-go: each action is weighted only by the rewards that come after it. Earlier rewards don't depend on that action, so dropping them cuts variance without adding bias. Number Search pays once at the end, so every action's reward-to-go equals R(τ).
GRPO: run the same task several times and use the mean as the baseline
The last section covers a simplified GRPO (Group Relative Policy Optimization), from DeepSeekMath (Shao et al.). As Fried tells it, most earlier work applying RL to language models trained a separate model to predict the baseline, which mattered a lot but was expensive. DeepSeek cared about efficiency and found a way to avoid the extra model.
The idea is simple: run the same task instance G times (resetting the environment each time), collect G rewards, and use their mean as the baseline. A baseline is supposed to be the policy's expected reward on this task, and the most direct way to estimate that is to actually run the policy a few times and average. It's an elegant idea. The original GRPO also divides by the group's standard deviation.
With the hidden number 11 and rewards 1, 1, 0, 0: the mean is 0.5, the standard deviation 0.5, so the advantages are +1, +1, −1, −1. The failures finally get a negative weight.
A batch holds several task instances, and each one is compared only with its own group:
| Task instance | Rewards of 4 rollouts | Group mean | Advantages |
|---|---|---|---|
| hidden 11 | 1, 1, 0, 0 | 0.50 | +1, +1, −1, −1 |
| hidden 3 | 1, 0, 0, 0 | 0.25 | +1.73, −0.58, −0.58, −0.58 |
| hidden 7 | 1, 1, 1, 0 | 0.75 | +0.58, +0.58, +0.58, −1.73 |
Look at the hidden-3 row: the single success gets a large positive weight. In the hidden-7 row, the single failure gets a heavy penalty. That matches the earlier intuition: a success on a hard task and a slip on an easy one both deserve big updates.
Fried said in class that the group size G is usually around 8, and that Apurva Gandhi will say more about choosing it in Lecture 12. That figure is the speaker's rule of thumb and is not on the slides; for comparison, DeepSeekMath's original setup samples 64 outputs per question, while the DrGRPO experiments use 8 responses per question, so the number varies a lot by setup. Asked whether this is unbiased, Fried said yes: the baseline comes only from rollouts on the same prompt and doesn't depend on the current action.
Fried was also explicit that the slides do not show the full GRPO loss from the paper. The paper adds an importance ratio, clipping (so no single step goes too far), and a KL term (so the policy doesn't drift too far from where it started). Those wait for Lecture 11.
Mechanism: group-relative advantages and the GRPO loss
For G rollouts on the same task instance:
$$\bar R = \frac{1}{G} \sum_{j=1}^{G} R(\tau_j), \qquad \sigma_R = \operatorname{std}{R(\tau_1), \ldots, R(\tau_G)}, \qquad \hat A_j = \frac{R(\tau_j) - \bar R}{\sigma_R}$$
For every $h_t$ in the group, $b(h_t) = \bar R$.
The policy-gradient loss (i indexes task instances in the batch):
$$-\sum_i R(\tau_i) \sum_t \log \pi_\theta(a_{i,t} \mid h_{i,t})$$
With GRPO's weights (j indexes rollouts in a group):
$$-\sum_i \frac{1}{G} \sum_{j=1}^{G} \hat A_{ij} \sum_t \log \pi_\theta(a_{ij,t} \mid h_{ij,t})$$
This is written for current-policy trajectories only, without the paper's ratio, clipping, or KL terms.
DrGRPO: drop two divisions
DrGRPO (Liu et al., in a paper titled Understanding R1-Zero-Like Training) points out that two of GRPO's divisions introduce bias. Assignment 3 asks you to implement both changes:
| GRPO | DrGRPO | Why | |
|---|---|---|---|
| Advantage | subtract the mean, divide by group std σ_R | subtract the mean only | Tasks the policy almost always passes or almost always fails have small σ_R, so dividing inflates their weight; standardizing also erases differences in difficulty |
| Response aggregation | divide each response by its own length |y_j| | divide by a global constant C | With varying lengths, a long wrong answer is penalized less per token than a short one, so the policy drifts toward long wrong answers |
C is the same for every response; the paper uses the generation budget. The rewards themselves don't change; what changes is the relative weight of tasks and responses. Fried mentioned in class (without a source) that many implementations calling themselves GRPO have already dropped the standard-deviation division.
Mechanism: DrGRPO's per-token coefficients
For one task instance, every generated token in trajectory j shares that trajectory's advantage, divided by the constant C:
$$L_{\text{DrGRPO}} = -\frac{1}{G} \sum_{j=1}^{G} \frac{1}{C} \sum_{t : m_{j,t} = 1} \hat A_j \log \pi_\theta(u_{j,t} \mid u_{j,<t}), \qquad \hat A_j = R(\tau_j) - \bar R$$
For example, each token in τ₁ (success) gets +0.5 / C, and each token in τ₃ (failure) gets −0.5 / C. The paper's full objective also has a ratio and clipping, omitted here.
All-pass and all-fail groups give no gradient
When every reward in a group is the same, subtracting the mean leaves all zeros, and the group contributes no gradient:
- [0, 0, 0, 0]: possibly too hard for the current policy.
- [1, 1, 1, 1]: possibly too easy.
If every trajectory fails, you can choose different task instances, broaden exploration, add hints to the prompt so the model has a better chance of succeeding, add demonstrations (off-policy RL can put correct demonstrations into the group; more simply, fine-tune on demonstrations before RL), or provide a more informative reward, known as reward shaping, such as rewarding the model for finding the right file.
How GRPO relates to PPO
GRPO was proposed as a more efficient alternative to PPO. PPO trains a separate value model to estimate the baseline; GRPO replaces it with group-relative advantages. Accounts of what that value model looks like differ: Fried said in class that it usually shares a backbone with the policy, while DeepSeekMath §4.1.1 describes it as "typically another model of comparable size as the policy model," which is its stated motivation for GRPO; InstructGPT's PPO, for instance, used a separate 6B value model. PPO's ratios, clipping, and reference-model KL term are covered in Lecture 11.
Back to the model: four methods, one weight
Fried's summary slide collapses the whole lecture into one formula:
$$\nabla_\theta J = \sum_t w_t, \nabla_\theta \log \pi_\theta(a_t \mid h_t)$$
| Method | Weight w_t |
|---|---|
| ReST-EM | 1 on successful trajectories, 0 on the rest (SFT on what was kept) |
| REINFORCE | R(τ), the reward of the trajectory the action came from |
| GRPO / DrGRPO | Â_j, the reward relative to other rollouts on the same task |
Back to the three gaps from the start:
- Task mismatch: the objective is expected reward, so any successful trajectory is reinforced.
- Data mismatch: with the right baseline, zero-reward trajectories also teach something.
- Exposure bias: trajectories come from the current policy. Fried likens it to the model driving the car during training, so it has to learn to recover from its own mistakes.
The price is trying things over and over in an environment that may be expensive, and that adds noise to the gradients. This lecture handled part of that; Lectures 11 and 12 handle more.
What this lecture changes for agent engineers
- Check your loss mask first. If you write your own RL training, tool outputs and prompt tokens must have m=0. With the wrong mask, the model is trained to predict the environment instead of learning to make decisions.
- Check that group rewards vary. Before running GRPO, sample 8 rollouts on a handful of tasks and look at the reward distribution. If most tasks are all 0 or all 1, adjust task difficulty or warm up with SFT before you start training.
- Try ReST first with binary rewards. It reuses your SFT pipeline unchanged and is the cheapest way to check whether the reward signal in your environment is learnable at all.
Going deeper
Readings listed for this lecture on the official schedule:
- ReST: Reinforced Self-Training for Language Modeling (Gulcehre et al., 2023)
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (ReST-EM, Singh et al., 2023)
- DeepSeekMath (Shao et al., 2024); the GRPO objective is in §4.1.1 and the group-normalized advantage in §4.1.2
- Understanding R1-Zero-Like Training: A Critical Perspective (DrGRPO, Liu et al., 2025)
The course schedule also lists Sean Welleck's RL Fundamentals slides from CMU ANLP; slide 7 of this lecture (how RL closes SFT's three gaps) is adapted from them.
Further reading on this site (the same algorithms from other angles; not a substitute for this lecture):
- CS336 Lecture 15: SFT Teaches Imitation; RLHF Begins Direct Preference Optimization
- CS336 Lecture 16: RLVR Scales Reasoning with Verifiable Rewards, but GRPO Is Not Free PPO
- Deep Reinforcement Learning: Putting RLHF Back Inside the RL Frame (CS230)
- CME295 Lecture 5: RLHF and DPO Add the Negative Signal
- CME295 Lecture 6: How Reasoning Models Learn to Think Longer, and What GRPO Drops from PPO
References
- CMU 11-768 AI Agents course website
- Lecture 9 slides: RL Foundations for Agents
- Lecture 9 recording
- Williams, 1992. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning (Machine Learning 8, 229–256; public copy on the UMass CS687 course page. REINFORCE and the reinforcement baseline are defined in §4, episodic REINFORCE in §5; the derivation in this post follows the slides, whose notation differs from the paper)
- Gulcehre et al., 2023. Reinforced Self-Training (ReST) for Language Modeling
- Singh et al., 2023. Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- Shao et al., 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Ouyang et al., 2022. Training language models to follow instructions with human feedback (Appendix C.4, PPO's separate value model)
- Liu et al., 2025. Understanding R1-Zero-Like Training: A Critical Perspective
- DeepSeek-AI, 2025. DeepSeek-R1
- Lewis et al., 2017. Deal or No Deal? End-to-End Learning for Negotiation Dialogues
- Sean Welleck, CMU ANLP Spring 2026. RL Fundamentals
Loading...