Skip to content

CS234 Learning from Demonstrations: Behavioral Cloning, DAgger, Inverse RL, and MaxEnt IRL

Sep 30, 20261 min
TL;DRWhen you have expert demonstrations but no reward, the second half of CS234 L7 offers three routes. Behavioral cloning copies actions with supervised learning. DAgger fixes its compounding errors by querying the expert along the learner's own path. Inverse RL instead infers what reward the expert is optimizing. Inferring rewards runs into the fact that infinitely many rewards explain the same demonstrations; feature matching and the maximum-entropy principle are two ways to pin down an answer. This material sets up the next post on RLHF: swap demonstrations for preferences and the problem keeps almost the same shape.

🌏 中文版

Edition note: this guide follows the CS234 Winter 2026 slides: Lecture 7 pp.25–62 and Lecture 8 pp.6–17 (PDF page numbers). The 2026 recordings are for enrolled students only. The public recordings are videos 7 and 8 of the Spring 2024 playlist, used here only as a listening supplement, with timestamps taken from the YouTube chapter markers. Access level A3 (defined in the global AI/CS course map). Every fact was checked on 2026-09-30 against those PDFs and video pages.

Series: previous A2: implementing REINFORCE, a baseline, and PPO | next Learning from human preferences: Bradley-Terry, RLHF, DPO | Series overview

Until now the reward has always been given. The second half of L7 asks a different question. The world already has very good decision-makers: drivers, pilots, doctors. Can we learn directly from their demonstrations without writing a reward first?

The slides give the motivation in one line. Having humans provide a reward signal while the RL algorithm acts is cheap supervision, but its sample complexity is high. The alternative is imitation learning.

Slide ranges and the 2024 videos

The 2026 PDF boundaries do not match the topics, so this post cuts by topic:

2026 slidesContentSpring 2024 video (chapter times)
L7 pp.26–30Learning from past decisions, reward shaping, problem setupVideo 07 from 45:26, "Introduction to imitation learning"
L7 pp.31–39Behavioral cloning, ALVINN, compounding errors, DAggerVideo 07 50:03–1:03:48
L7 pp.40–50Reward learning, linear-feature IRL, feature matching, ambiguityVideo 07 from 1:03:48
L7 pp.51–62MaxEnt IRL, from IRL to policies, summaryVideo 08 from 4:28 to 52:26
L8 pp.6–17"How Can RL Enable Transformative LLM?", DAgger and feature-reward recap, Imitation Learning SummaryVideo 08 has a matching overview at 1:11–4:28

Video 08 is titled "Offline RL 1" in the playlist, but its YouTube chapters show MaxEnt IRL and the start of RLHF. The 2026 slides have no dedicated offline RL lecture, and this post does not invent 2026 content for it.

Why learn from demonstrations

Reward shaping is the first motivation. Rewards that are dense in time guide an agent closely, but who supplies them? The slides list two options. You can design them by hand, which is often brittle. Or you can specify them implicitly through demonstrations. Examples include simulated highway driving (Abbeel & Ng 2004 and others) and parking-lot navigation (Abbeel, Dolgov, Ng, and Thrun, IROS 2008).

Imitation learning helps when it is easier for an expert to demonstrate the behavior than to write a reward that produces it, or to write the policy directly.

The problem setup (L7 p.30):

  • A known state space, action space, and transition model P(s′|s, a)
  • No reward function R
  • One or more expert demonstrations (s₀, a₀, s₁, a₁, …), with actions drawn from the expert policy π*

Three questions branch off from there:

  1. Behavioral cloning: can we learn the expert policy directly with supervised learning?
  2. Inverse RL: can we recover R?
  3. Apprenticeship learning via inverse RL: can we use the recovered R to produce a good policy?

Behavioral cloning: RL as supervised learning

The recipe is direct. Fix a policy class, such as a neural network or a decision tree, and fit it on (s₀, a₀), (s₁, a₁), and so on. The slides name two early successes: Pomerleau's ALVINN at NIPS 1989, which learned to drive from images, and Sammut et al. at ICML 1992, who learned to fly in a flight simulator.

The slides also stress that it often works very well in practice, especially with BC-RNN. They cite Mandlekar et al., "What Matters in Learning from Offline Human Demonstrations for Robot Manipulation" (CoRL 2021). The verdict: "Extensively used in practice."

The catch: compounding errors

Supervised learning assumes i.i.d. (s, a) pairs and ignores temporal structure. If errors were independent in time with probability ε per step, expected total errors would be about εT.

In an MDP, training and test distributions differ:

  • In training, sₜ comes from the distribution induced by the expert π*
  • At test time, sₜ comes from the distribution induced by the learned policy π_θ

One mistake puts the agent in states the expert never visited, and later errors become more likely. The slides' rough intuition is E[total errors] ≲ ε(T + (T−1) + … + 1) ≈ εT². For the rigorous result they point to Theorem 2.1 of Ross & Bagnell, AISTATS 2010.

DAgger: ask the expert where you actually go

The idea from Ross, Gordon, and Bagnell 2011: collect more expert labels along the path taken by the behavior-cloned policy. The algorithm on L7 p.39:

  1. Start with an empty dataset D and any π̂₁
  2. In round i, let πᵢ = βᵢπ* + (1−βᵢ)π̂ᵢ and run it for T steps
  3. Ask the expert for π*(s) at every visited state to get Dᵢ
  4. Set D ← D ∪ Dᵢ and train π̂ᵢ₊₁ on D
  5. Return the best π̂ᵢ on validation

The slides say this yields a stationary deterministic policy that performs well under its own induced state distribution. Then they ask: "Key limitation?" Look at step 3 for the answer. Every round needs an expert on hand to label arbitrary states on demand.

Inverse RL: what is the expert optimizing?

Flip the question. If the expert's policy is optimal, what can we infer about R?

The quiz on L7 pp.42–43 gives the answer: infinitely many R make the expert's policy optimal. This ambiguity is the central difficulty of IRL.

Linear-feature rewards

Restrict to R(s) = wᵀx(s), where x is a state-feature vector and w the weights to learn. Plug it into the value function:

V^π(s₀) = E[Σ γᵗ wᵀx(sₜ)] = wᵀ E[Σ γᵗ x(sₜ)] = wᵀμ(π)

μ(π) is the discounted weighted feature frequency under π. This mirrors linear value-function approximation, except that now the reward is the linear part.

If the demonstrations come from an optimal policy, finding w means finding a w* with wᵀμ(π) ≥ wᵀμ(π) for every π ≠ π.

Feature matching

Abbeel & Ng (2004) observed that a policy π is guaranteed to do as well as the expert if its discounted feature expectations are close enough to the expert's. Precisely: if ‖μ(π) − μ(π*)‖₁ ≤ ε, then for every w with ‖w‖∞ ≤ 1, |wᵀμ(π) − wᵀμ(π*)| ≤ ε, by Hölder's inequality.

The result is elegant. You do not need the true w; if the features match, the performance matches.

But the ambiguity does not go away; it gains a layer. L7 p.49 notes that infinitely many rewards share the same optimal policy, and infinitely many stochastic policies can match the feature counts. Which one should you pick? The slides point to two key papers: MaxEnt IRL by Ziebart et al. (AAAI 2008) and GAIL by Ho & Ermon (NeurIPS 2016). The lecture continues with the first.

MaxEnt IRL: among all matching distributions, pick the least opinionated

Keep R(s) = wᵀx(s). This time the feature count is summed over a single trajectory, μ_τ = Σ x(sᵢ). The slides flag that this differs slightly from the earlier definition. The average over m demonstrations is μ̃.

In a deterministic MDP with a linear reward, a policy is fully specified by its distribution over H-step trajectories. So the question becomes: given m demonstrations, which trajectory distribution should we choose?

The principle of maximum entropy: add no preference beyond matching the demonstrations' feature expectations. As an optimization, maximize −Σ P(τ) log P(τ) subject to Σ P(τ)μ_τ = μ̃ and Σ P(τ) = 1.

With a linear reward, this is equivalent to maximizing the likelihood of the demonstrations under an exponential-family distribution:

P(τⱼ | w) = exp(wᵀμ_τⱼ) / Z(w)

The slides' reading: a strong preference for low-cost paths, with equal-cost paths equally likely. Stochastic MDPs multiply in the transition probabilities along the trajectory.

Learning w

Choose w to maximize the log-likelihood of the demonstrations. The gradient is a clean difference:

∇L(w) = μ̃ − Σ_τ P(τ | w)μ_τ = μ̃ − Σ_s D(s)x(s)

That is, the demonstrations' feature counts minus the learner's expected feature counts under the current reward. The second term can be written with state-visitation frequencies D(s). L7 p.58 gives the algorithm for D(s): a backward pass computes local action probabilities, a forward pass propagates state-visitation frequencies, and a final step sums over time.

The slides then ask whether computing this requires the transition model. L7 p.59 answers that the original formulation needs the transition model, or the ability to act in the world and sample transitions. Then comes a second question: did behavioral cloning need it? Keep this contrast in mind. BC needs only demonstrations; IRL needs demonstrations plus a model or interaction.

The slides call the maximum-entropy approach "hugely influential": it offers a principled way to choose among the many possible rewards.

From IRL back to policies

Once you have a reward, what you usually want is a policy that matches or beats the expert. One approach on L7 p.60 is to feed the learned reward into ordinary RL. Then the slides ask whether we can learn the desired policy more directly. The lecture leaves that open, but that question is where the GAIL line of work begins.

Summary and the bridge onward

Three points from the L7 p.61 summary are worth keeping:

  • Imitation learning can greatly reduce the data needed to learn a good policy
  • Combining inverse RL / learning from demonstration with online RL is an active direction
  • Often we only have preference pairs (y₁ ≻ y₂), not demonstrations. The slides call this the "dueling bandits" setting and say it will return shortly and in Assignment 3

The opening of L8 compresses this into an "Imitation Learning Summary" slide: very powerful, many extensions, and maximum-entropy reward learning is an important idea. Before that, L8 p.6 shows a ChatGPT screenshot answering "write a program to demonstrate how RLHF works," as the bridge from robot demonstrations to LLMs.

The next post picks up there. Replace "expert demonstrations" with "a human says A is better than B," and the IRL problem turns almost unchanged into RLHF reward modeling.

How to self-study it

  1. Read L7 pp.26–39 and stop at DAgger's "Key limitation?" Write your answer down before moving on.
  2. For pp.40–50, copy the three-line derivation of V^π = wᵀμ(π) onto paper. The Hölder step in feature matching is one line.
  3. Pair the MaxEnt section (pp.51–59) with the chapters from "Max entropy IRL math" through "Max entropy IRL algorithm" in 2024 video 08. The video walks through the gradient derivation in more detail than the slides.
  4. For hands-on practice: CS234's assignments have no imitation-learning question, but CS224R HW1 has you implement BC and DAgger.

One thing to do tonight: make a table of what BC, DAgger, and IRL each need. Do they need the transition model? An expert on call? Interaction with the environment? That table is the answer key for every "Check your understanding" in L7.

Further reading

References