🌏 中文版
Source term: Based on the Spring 2026 08_cs224r_reward_learning_2026 slides (scheduled 2026-04-24). The companion video is the Spring 2025 Lecture 8 recording (supplement). The title matches, but the opening offline RL recap differs: the 2025 Lecture 8 slides are titled "Conservative Offline RL and Reward Learning" and recap conservative methods, while the 2026 version recaps Lecture 7's two key ideas and adds a π*0.6 example. The three reward-learning subsections are the same in both years. This is post 10 in the Reading Stanford CS224R series.
Since Lecture 1, CS224R has treated the reward r(s, a) as given. Lecture 8 finally asks who supplies that number.
The slides give two learning goals:
- why task specification is hard (and why naive methods fail)
- methods for learning rewards from human supervision
The plan has two parts. Part one is an offline RL recap and example, marked "Part of HW3." Part two is reward learning, where the preference subsection is marked "Part of default project" and "How LLMs are supervised!"
Wrapping up offline RL
This section compresses Lecture 7 into two key ideas. The setting is unchanged: data from an unknown πβ, rewards to maximize under πθ.
The slide first adds a contrast Lecture 7 left implicit. Querying the Q-function on OOD actions causes overestimation. In online RL, data from the new policy corrects those errors in later iterations. Offline RL has no additional data, so it has to be more conservative.
- Key idea 1: train the policy only on actions sampled from the dataset, for example with advantage-weighted regression. AWR also fits the value of πβ rather than πθ, so even the value function never queries OOD actions.
- Key idea 2: use an asymmetric expectile loss to fit the value of a policy better than πβ, then fit Q from that V. This is what IQL does.
Then comes a real example. π*0.6 (2025) uses an "iterated offline RL" recipe for robot post-training, in three steps:
- Collect a large batch of data with π (a mix of roll-outs and DAgger).
- Fit V^πβ with Monte Carlo.
- Train an advantage-conditioned policy π(aₜ | sₜ, Â(sₜ, aₜ)).
The slide goes no further than this; see the paper for details. The example ties together DAgger from Lecture 2, Monte Carlo values, and the advantage from Lecture 7.
Where do rewards come from?
On the left, the slide shows computer games (citing Mnih et al. 2015 on Atari), which come with a score. On the right is the real world: robotics, dialogue, autonomous driving, labeled only "what is the reward?" In practice people often use a proxy.
Is there an easier way to supervise a task? The course has already covered one: directly imitating an expert's actions. The slide lists its limits:
- it doesn't reason about outcomes or dynamics
- the expert may have different degrees of freedom than the robot
- some tasks can't be demonstrated at all
So the question becomes: can we reason about what the expert is trying to achieve?
Route 1: learn rewards from goal examples
Goal classifiers
The most direct idea is to learn a classifier that tells goal states apart from other states. The slide's example task is "put the pencil case behind the notebook":
- Collect examples of successful and unsuccessful states (inside and outside the goal set G).
- Train a binary classifier with inputs sᵢ and labels 1(sᵢ ∈ G).
- Run RL with the classifier's output as the reward.
What can go wrong? RL seeks out states the classifier thinks are good, and those may be states the classifier was never trained on. The policy finds the classifier's weaknesses, not a solution to the task.
Add visited states as negatives
The fix on the slide comes from Fu et al. 2018 (VICE):
- Collect an initial set of successful states D+ and unsuccessful states D−.
- Update the classifier with D+ and D−, balancing each batch 50/50.
- Collect experience with policy π.
- Update π with the classifier-based reward.
- Add visited states to the negatives: D− ← D− ∪ {sₜ}.
The class first discusses three questions: will the classifier be accurate, will the policy work, and what will the classifier output for successful states? The slide's answers:
- The classifier can't be exploited, because anywhere the policy goes becomes a negative.
- But what if some visited states are actually successful?
- As long as batches stay balanced, the classifier still outputs p ≥ 0.5 for successful states.
Results on a robot
The slide cites Sharma et al. 2023. They collected 50 demonstrations, used the final states as success examples, and seeded the RL replay buffer with the demos. Directly imitating the demos reached a 26% success rate. An RL policy trained with the learned classifier reached 62%. The slide adds a note: regularizing the classifier matters.
This is how GANs work
A side note on the slide makes the connection: GANs follow the same recipe. Train a classifier to tell real data from generated data, then train a generator to produce data the classifier thinks is real. At convergence, the generator matches the data distribution p(x). The examples are ViT-VQGAN and Phenaki video generation.
Map it over: the policy is the generator, success examples are the real data, and the goal classifier is the discriminator.
The slide's summary of this route:
| Pro | A practical framework for task specification |
| Caveat | Adversarial training can be unstable (the GAN literature has many regularization tricks) |
| Con | Requires examples of desired behavior or outcomes |
One more point from that summary: with success examples you can learn a goal classifier; with full demos you can learn a full reward.
Route 2: learn rewards from human preferences
Comparing is easier than scoring
What if you skip demos and goal examples and ask people for feedback on policy roll-outs instead? The slide lists two ways to ask:
- "How good is this trajectory?"
- "Which trajectory is better?"
Its verdict: relative preferences are easier to provide.
From preferences to a reward function
A person says τw is better than τl, written τw ≻ τl. We want a reward rθ whose sum over τw exceeds its sum over τl. τ can be a full or partial roll-out.
The slide frames it this way: humans are classifying which trajectory is better, so the reward should be discriminative too. Concretely, define σ(rθ(τa) − rθ(τb)) as the estimated probability that τa ≻ τb, then maximize the log probability:
max_θ E_{τw, τl} [ log σ( rθ(τw) − rθ(τl) ) ]
The complete algorithm:
- From a dataset {τᵢ}, sample batches of k trajectories and ask humans to rank them. (For LLMs, all k share the same prompt.)
- Compute rθ for each trajectory under the current reward model.
- For all k-choose-2 pairs per batch, compute the gradient of the objective above.
- Update θ with that gradient.
The slide notes that this can run inside the loop of online RL.
Two examples
- Christiano et al. 2017 learn rewards inside the online RL loop; the slide says they used 900 human preference queries.
- Sadigh et al. 2017 (RSS) learn a driving reward from preferences to weight different factors.
Applied to LLMs: RLHF
For LLMs: given a prompt x, sample two replies y and y′, ask a human which is better, and train a reward model r(x, y) that judges how good reply y is for prompt x.
The slide places this in a three-stage LLM training pipeline:
- Large-scale pretraining: next-token prediction on mixed-quality data.
- Supervised fine-tuning on higher-quality (prompt, response) pairs.
- RLHF:
- 3a. Gather preference data.
- 3b. Train the reward model.
- 3c. Run RL to maximize the reward (e.g. with PPO).
RLAIF: swap the human for an AI
The last variation replaces the human with another language model, asking it "which of these responses is less harmful?" The source is Anthropic's Constitutional AI (2022). The slide's key insight: critique is easier than generation.
Summary: the tradeoff between the two routes
The first line of the summary slide: rewards can't be taken for granted.
| Learning from goals and demos | Learning from human preferences | |
|---|---|---|
| Pros | A practical framework for task specification | Pairwise preferences are easy to give, with no goal examples or demos needed; deployed at scale |
| Caveat | Adversarial training can be unstable | |
| Cons | Requires examples of desired behavior or outcomes | May require supervision in the RL loop, which usually takes more human time |
The slide leaves a thought exercise: what other forms of feedback or supervision might help?
The last content slide points to the whole area of "unsupervised" RL: can agents propose their own goals? The example is asymmetric self-play (Sukhbaatar et al. 2018), framed as a two-player game between a goal-setter and a goal-reacher.
Something to try tonight
Pick an agent or LLM feature you work on and write down what its "reward" is today: human ratings, rules, an LLM judge, or user thumbs-up. Then check it against this lecture's two questions. When a policy optimizes against it, could it find outputs that score high without doing the task? If so, could you do what VICE does and keep feeding the exploited outputs back in as negatives?
Further reading
- CS224R Lecture 9: RLHF and Preference Optimization: the next lecture in this series, from reward model + RL to DPO
- CME295: Preference Tuning: RLHF and DPO from the LLM side
- CS336: SFT and RLHF: the post-training pipeline from the implementation side
Series navigation: Previous: Lecture 7: Offline RL | Next: HW3: AWAC, IQL, and Stitching on AntMaze | Series overview
References
- CS224R course homepage and schedule (Spring 2026)
- Lecture 8 slides: Recap of offline RL + Reward Learning (2026)
- Spring 2025 Lecture 8: Reward Learning (YouTube, supplement)
- Christiano et al. Deep Reinforcement Learning from Human Preferences (arXiv 1706.03741)
- Fu et al. Variational Inverse Control with Events (arXiv 1805.11686)
- Sharma et al. Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning (arXiv 2303.01488)
- Bai et al. Constitutional AI: Harmlessness from AI Feedback (arXiv 2212.08073)
- π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)
- Sukhbaatar et al. Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play (arXiv 1703.05407)
Loading...