🌏 中文版
Edition note: this guide follows the CS234 Winter 2026 slides: Lecture 7 pp.25–62 and Lecture 8 pp.6–17 (PDF page numbers). The 2026 recordings are for enrolled students only. The public recordings are videos 7 and 8 of the Spring 2024 playlist, used here only as a listening supplement, with timestamps taken from the YouTube chapter markers. Access level A3 (defined in the global AI/CS course map). Every fact was checked on 2026-09-30 against those PDFs and video pages.
Series: previous A2: implementing REINFORCE, a baseline, and PPO | next Learning from human preferences: Bradley-Terry, RLHF, DPO | Series overview
Until now the reward has always been given. The second half of L7 asks a different question. The world already has very good decision-makers: drivers, pilots, doctors. Can we learn directly from their demonstrations without writing a reward first?
The slides give the motivation in one line. Having humans provide a reward signal while the RL algorithm acts is cheap supervision, but its sample complexity is high. The alternative is imitation learning.
Slide ranges and the 2024 videos
The 2026 PDF boundaries do not match the topics, so this post cuts by topic:
| 2026 slides | Content | Spring 2024 video (chapter times) |
|---|---|---|
| L7 pp.26–30 | Learning from past decisions, reward shaping, problem setup | Video 07 from 45:26, "Introduction to imitation learning" |
| L7 pp.31–39 | Behavioral cloning, ALVINN, compounding errors, DAgger | Video 07 50:03–1:03:48 |
| L7 pp.40–50 | Reward learning, linear-feature IRL, feature matching, ambiguity | Video 07 from 1:03:48 |
| L7 pp.51–62 | MaxEnt IRL, from IRL to policies, summary | Video 08 from 4:28 to 52:26 |
| L8 pp.6–17 | "How Can RL Enable Transformative LLM?", DAgger and feature-reward recap, Imitation Learning Summary | Video 08 has a matching overview at 1:11–4:28 |
Video 08 is titled "Offline RL 1" in the playlist, but its YouTube chapters show MaxEnt IRL and the start of RLHF. The 2026 slides have no dedicated offline RL lecture, and this post does not invent 2026 content for it.
Why learn from demonstrations
Reward shaping is the first motivation. Rewards that are dense in time guide an agent closely, but who supplies them? The slides list two options. You can design them by hand, which is often brittle. Or you can specify them implicitly through demonstrations. Examples include simulated highway driving (Abbeel & Ng 2004 and others) and parking-lot navigation (Abbeel, Dolgov, Ng, and Thrun, IROS 2008).
Imitation learning helps when it is easier for an expert to demonstrate the behavior than to write a reward that produces it, or to write the policy directly.
The problem setup (L7 p.30):
- A known state space, action space, and transition model P(s′|s, a)
- No reward function R
- One or more expert demonstrations (s₀, a₀, s₁, a₁, …), with actions drawn from the expert policy π*
Three questions branch off from there:
- Behavioral cloning: can we learn the expert policy directly with supervised learning?
- Inverse RL: can we recover R?
- Apprenticeship learning via inverse RL: can we use the recovered R to produce a good policy?
Behavioral cloning: RL as supervised learning
The recipe is direct. Fix a policy class, such as a neural network or a decision tree, and fit it on (s₀, a₀), (s₁, a₁), and so on. The slides name two early successes: Pomerleau's ALVINN at NIPS 1989, which learned to drive from images, and Sammut et al. at ICML 1992, who learned to fly in a flight simulator.
The slides also stress that it often works very well in practice, especially with BC-RNN. They cite Mandlekar et al., "What Matters in Learning from Offline Human Demonstrations for Robot Manipulation" (CoRL 2021). The verdict: "Extensively used in practice."
The catch: compounding errors
Supervised learning assumes i.i.d. (s, a) pairs and ignores temporal structure. If errors were independent in time with probability ε per step, expected total errors would be about εT.
In an MDP, training and test distributions differ:
- In training, sₜ comes from the distribution induced by the expert π*
- At test time, sₜ comes from the distribution induced by the learned policy π_θ
One mistake puts the agent in states the expert never visited, and later errors become more likely. The slides' rough intuition is E[total errors] ≲ ε(T + (T−1) + … + 1) ≈ εT². For the rigorous result they point to Theorem 2.1 of Ross & Bagnell, AISTATS 2010.
DAgger: ask the expert where you actually go
The idea from Ross, Gordon, and Bagnell 2011: collect more expert labels along the path taken by the behavior-cloned policy. The algorithm on L7 p.39:
- Start with an empty dataset D and any π̂₁
- In round i, let πᵢ = βᵢπ* + (1−βᵢ)π̂ᵢ and run it for T steps
- Ask the expert for π*(s) at every visited state to get Dᵢ
- Set D ← D ∪ Dᵢ and train π̂ᵢ₊₁ on D
- Return the best π̂ᵢ on validation
The slides say this yields a stationary deterministic policy that performs well under its own induced state distribution. Then they ask: "Key limitation?" Look at step 3 for the answer. Every round needs an expert on hand to label arbitrary states on demand.
Inverse RL: what is the expert optimizing?
Flip the question. If the expert's policy is optimal, what can we infer about R?
The quiz on L7 pp.42–43 gives the answer: infinitely many R make the expert's policy optimal. This ambiguity is the central difficulty of IRL.
Linear-feature rewards
Restrict to R(s) = wᵀx(s), where x is a state-feature vector and w the weights to learn. Plug it into the value function:
V^π(s₀) = E[Σ γᵗ wᵀx(sₜ)] = wᵀ E[Σ γᵗ x(sₜ)] = wᵀμ(π)
μ(π) is the discounted weighted feature frequency under π. This mirrors linear value-function approximation, except that now the reward is the linear part.
If the demonstrations come from an optimal policy, finding w means finding a w* with wᵀμ(π) ≥ wᵀμ(π) for every π ≠ π.
Feature matching
Abbeel & Ng (2004) observed that a policy π is guaranteed to do as well as the expert if its discounted feature expectations are close enough to the expert's. Precisely: if ‖μ(π) − μ(π*)‖₁ ≤ ε, then for every w with ‖w‖∞ ≤ 1, |wᵀμ(π) − wᵀμ(π*)| ≤ ε, by Hölder's inequality.
The result is elegant. You do not need the true w; if the features match, the performance matches.
But the ambiguity does not go away; it gains a layer. L7 p.49 notes that infinitely many rewards share the same optimal policy, and infinitely many stochastic policies can match the feature counts. Which one should you pick? The slides point to two key papers: MaxEnt IRL by Ziebart et al. (AAAI 2008) and GAIL by Ho & Ermon (NeurIPS 2016). The lecture continues with the first.
MaxEnt IRL: among all matching distributions, pick the least opinionated
Keep R(s) = wᵀx(s). This time the feature count is summed over a single trajectory, μ_τ = Σ x(sᵢ). The slides flag that this differs slightly from the earlier definition. The average over m demonstrations is μ̃.
In a deterministic MDP with a linear reward, a policy is fully specified by its distribution over H-step trajectories. So the question becomes: given m demonstrations, which trajectory distribution should we choose?
The principle of maximum entropy: add no preference beyond matching the demonstrations' feature expectations. As an optimization, maximize −Σ P(τ) log P(τ) subject to Σ P(τ)μ_τ = μ̃ and Σ P(τ) = 1.
With a linear reward, this is equivalent to maximizing the likelihood of the demonstrations under an exponential-family distribution:
P(τⱼ | w) = exp(wᵀμ_τⱼ) / Z(w)
The slides' reading: a strong preference for low-cost paths, with equal-cost paths equally likely. Stochastic MDPs multiply in the transition probabilities along the trajectory.
Learning w
Choose w to maximize the log-likelihood of the demonstrations. The gradient is a clean difference:
∇L(w) = μ̃ − Σ_τ P(τ | w)μ_τ = μ̃ − Σ_s D(s)x(s)
That is, the demonstrations' feature counts minus the learner's expected feature counts under the current reward. The second term can be written with state-visitation frequencies D(s). L7 p.58 gives the algorithm for D(s): a backward pass computes local action probabilities, a forward pass propagates state-visitation frequencies, and a final step sums over time.
The slides then ask whether computing this requires the transition model. L7 p.59 answers that the original formulation needs the transition model, or the ability to act in the world and sample transitions. Then comes a second question: did behavioral cloning need it? Keep this contrast in mind. BC needs only demonstrations; IRL needs demonstrations plus a model or interaction.
The slides call the maximum-entropy approach "hugely influential": it offers a principled way to choose among the many possible rewards.
From IRL back to policies
Once you have a reward, what you usually want is a policy that matches or beats the expert. One approach on L7 p.60 is to feed the learned reward into ordinary RL. Then the slides ask whether we can learn the desired policy more directly. The lecture leaves that open, but that question is where the GAIL line of work begins.
Summary and the bridge onward
Three points from the L7 p.61 summary are worth keeping:
- Imitation learning can greatly reduce the data needed to learn a good policy
- Combining inverse RL / learning from demonstration with online RL is an active direction
- Often we only have preference pairs (y₁ ≻ y₂), not demonstrations. The slides call this the "dueling bandits" setting and say it will return shortly and in Assignment 3
The opening of L8 compresses this into an "Imitation Learning Summary" slide: very powerful, many extensions, and maximum-entropy reward learning is an important idea. Before that, L8 p.6 shows a ChatGPT screenshot answering "write a program to demonstrate how RLHF works," as the bridge from robot demonstrations to LLMs.
The next post picks up there. Replace "expert demonstrations" with "a human says A is better than B," and the IRL problem turns almost unchanged into RLHF reward modeling.
How to self-study it
- Read L7 pp.26–39 and stop at DAgger's "Key limitation?" Write your answer down before moving on.
- For pp.40–50, copy the three-line derivation of V^π = wᵀμ(π) onto paper. The Hölder step in feature matching is one line.
- Pair the MaxEnt section (pp.51–59) with the chapters from "Max entropy IRL math" through "Max entropy IRL algorithm" in 2024 video 08. The video walks through the gradient derivation in more detail than the slides.
- For hands-on practice: CS234's assignments have no imitation-learning question, but CS224R HW1 has you implement BC and DAgger.
One thing to do tonight: make a table of what BC, DAgger, and IRL each need. Do they need the transition model? An expert on call? Interaction with the environment? That table is the answer key for every "Check your understanding" in L7.
Further reading
- The same topic in a deep RL course: CS224R L2: imitation learning and multimodal policies
- Another take on distribution shift: Berkeley CS285 L1–4: imitation learning, distribution shift, and RL basics
References
- CS234 Lecture 7 slides (post version, Winter 2026) — pp.25–62: reward shaping, BC, compounding errors, DAgger, linear IRL, feature matching, MaxEnt IRL
- CS234 Lecture 8 slides (post version, Winter 2026) — pp.6–17: the LLM bridge slide and the imitation learning summary
- CS234 course home page (Winter 2026) — schedule (Week 4, "Offline RL, Imitation Learning")
- Stanford CS234 Spring 2024 Lecture 7, "Policy Search 3" — second half covers imitation learning, DAgger, and IRL (per YouTube chapters)
- Stanford CS234 Spring 2024 Lecture 8, "Offline RL 1" — chapters show MaxEnt IRL and the start of RLHF
- Ross & Bagnell, Efficient Reductions for Imitation Learning (AISTATS 2010) — the compounding-error theorem cited in the slides
- Ross, Gordon & Bagnell, A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (2011) — DAgger
- Abbeel & Ng, Apprenticeship Learning via Inverse Reinforcement Learning (ICML 2004) — feature matching
- Ziebart et al., Maximum Entropy Inverse Reinforcement Learning (AAAI 2008) — MaxEnt IRL
- Ho & Ermon, Generative Adversarial Imitation Learning (NeurIPS 2016) — the other key paper named in the slides
Loading...