A reading of Stanford CS224R Spring 2026 through its 17 slide decks, three homeworks, and default project. It covers deep RL from imitation learning, policy gradients, and offline RL to RLHF, LLM reasoning, and robot VLAs, with the public Spring 2025 videos as a labeled supplement.
CS224R is Chelsea Finn's deep reinforcement learning course at Stanford. It runs from imitation learning to RL for LLMs and robot foundation models. For Spring 2026, all 17 slide decks, the three homework handouts with starter code, and the default project spec with starter code can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the 2026 recordings are Canvas-only, the midterm and its solutions are not public, and HW2 and HW3 require Modal. The public recordings are from Spring 2025, so this series uses them as a supplement and flags the differences lecture by lecture.
The first lecture of CS224R Spring 2026 does three things: covers logistics, explains why deep RL is worth learning, and turns 'behavior' into something you can learn. The core is a set of definitions (state, action, trajectory, reward, policy) and one objective: maximize expected total reward. It ends on an example: fit ℓ2 regression to drivers where some change lanes and some go straight, and the policy learns their average, a half lane change nobody demonstrated. That problem is where L2 starts.
Lecture 2 of CS224R Spring 2026 tackles two ways imitation learning fails. First, when demonstrations contain several reasonable behaviors, regression learns only their average. The fix is to make the policy a generative model (Gaussian mixtures, discretization plus autoregression, diffusion or flow matching) and to add action chunking. Second, compounding errors: once the policy slips, it reaches states the demonstrations never covered. The fix is DAgger or human-gated DAgger to collect corrections. The first two parts are exactly what HW1 covers.
Homework 1 of CS224R Spring 2026 tests imitation learning on a custom Flappy Bird environment. The policy predicts 20 future target heights at once and executes only the first 10. You implement MSE-regression behavior cloning, a flow matching policy, and DAgger, then compare them in easy and hard modes. The PDF, LaTeX template, and starter code are all public, and a CPU is enough to run it. Solutions, the autograder, and Gradescope are not public. This guide covers what each problem asks you to build and answer. It does not give solutions.
Policy gradient is the first online RL algorithm in CS224R. Its gradient looks almost exactly like the imitation learning gradient, except that each trajectory is weighted by its reward. Actions from good outcomes become more likely, and actions from bad outcomes become less likely. The raw version is very noisy, so L3 cuts the variance in two ways: count only future rewards (causality) and subtract the average reward (a baseline). It is also on-policy, so every gradient step needs fresh data. Importance sampling plus a KL constraint lets you take several steps on one batch.
Policy gradient can only judge good and bad from the rewards it actually received, so it wastes data. Actor-critic trains a second network, a value function (the critic), to estimate how good a state is, and uses it to compute advantages that weight the policy's (the actor's) gradient. There are three ways to estimate value: supervise directly with a rollout's summed rewards (Monte Carlo), supervise with this step's reward plus your own estimate of the next state (bootstrapping), or use an n-step return in between. L4 ends by pushing actor-critic off-policy, first by taking several gradient steps on one batch (where PPO starts) and then by reusing all past data from a replay buffer (where SAC starts).
PPO and SAC answer the same question: can you use an expensive batch of data more than once? PPO takes several gradient steps on one fresh batch and clips the new-to-old policy ratio to 1±ε. SAC keeps every past transition in a replay buffer and learns Q(s, a), so old data can still evaluate the current policy. PPO is stable and easy to tune. SAC is data-efficient and harder to tune.
Q-learning drops the actor from actor-critic: learn the optimal Q-function directly and act by taking the argmax. The price is that convergence is not guaranteed; even linear Q can diverge. Lecture 6 of CS224R pulls it back with three engineering tricks: a target network that holds the targets still, Double Q that separates choosing an action from valuing it to curb overestimation, and n-step returns that trade a little bias for speed.
CS224R Spring 2026 HW2 has three parts. First, tabular Q-learning on a 5×4 gridworld shows how reward design changes the learned path. Second, GAE plus PPO clipping tackles a hammer task that pays 1 only on completion. Third, an off-policy actor-critic with BC pretraining, a critic ensemble, and a higher UTD ratio, followed by a comparison of the two learning curves. The handout, starter code, and compute guide are public, but the assignment supports only Modal, and course credits go only to enrolled students.
Lecture 7 of CS224R (Spring 2026) asks how to learn a policy better than your data when all you have is a fixed dataset someone else collected. Running an off-policy algorithm like SAC on that data fails: the Q-function makes up values for actions the data never contains, and the policy goes looking for exactly those overestimated actions. The slides give two families of fixes. One trains the policy only on actions in the data (filtered BC, AWR, AWAC). The other uses an asymmetric expectile loss to estimate the value of a policy better than the data without ever querying out-of-data actions (IQL). Both can do something imitation learning can't: stitch good pieces of different trajectories together.
Lecture 8 of CS224R (Spring 2026) spends a few slides wrapping up offline RL, then asks the question the first seven lectures skipped: where does the reward come from? Games have scores. Real robots, dialogue, and driving usually don't. The slides offer two routes. The first trains a goal classifier on success examples and uses it as the reward, but RL learns to exploit the classifier's blind spots; the fix is to keep adding states the policy visits as negatives, the same structure as a GAN. The second asks people which of two trajectories is better and learns a reward with the Bradley-Terry-style objective log σ(r(τw) − r(τl)), the same method LLM RLHF uses. The lecture's number-one takeaway is one line: rewards can't be taken for granted.
HW3 in CS224R (Spring 2026) has you fill in two offline RL algorithms, AWAC and IQL, and compare them on D4RL's AntMaze. Problem 1 runs AWAC on antmaze-umaze and antmaze-medium-diverse. Problem 2 compares IQL expectiles ζ = 0.2 and 0.9, runs the better value on medium-diverse, and then tests whether IQL can stitch a better path out of a PointMass dataset whose best return is only −46, against a filtered BC baseline that keeps the top 10% of trajectories. The PDF, LaTeX template, and starter code are public, but the assignment is meant to run on Modal, and course credits go only to enrolled students. This post covers the tasks and setup only, with no solutions.
Lecture 9 of CS224R Spring 2026 is a guest lecture by Archit Sharma, with slides adapted from CS224N. The spine is one chain of reasoning. Instruction tuning can't handle tasks with no right answer or errors of unequal weight, so we optimize human preferences directly. Human ratings are expensive and noisy, so we collect pairwise comparisons and fit a Bradley-Terry reward model. RLHF uses that model as the reward and runs policy gradient with a KL penalty. DPO uses the closed-form solution of the KL-constrained problem to write the reward as a log-ratio of policies, which turns the whole thing into a binary classification loss. The last part covers the frontier: reward hacking, verifiable rewards, and AI feedback in place of human feedback.
Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.
The CS224R Spring 2026 default project has you implement three stages on Qwen2.5-0.5B Base for the Countdown arithmetic reasoning task: SFT warm-start, IPO preference optimization, and RLOO with a rule-based verifier reward. All three are compared with the same vLLM evaluation, followed by a research extension of your choice. For the implementation, high-level trainers like SFTTrainer are banned, and so is any AI tool assistance; only the extension is exempt. The extension is half the grade for this project, and it's graded on methodology and documentation, not score. The starter code and datasets are public. What outside readers lack is Modal credits and the autograder.
Model-based RL first learns a dynamics model that predicts s_{t+1}, then uses it in one of two ways: to generate extra training data (Dyna, MBPO) or to think a few steps ahead before acting (planning). The thread running through Lecture 11 of CS224R is how to avoid being dragged down by model error. Start synthetic rollouts from real states and keep them short, average errors out with an ensemble of models, and attach a value function to the tail of long-horizon plans. Whether a model is worth learning depends on whether it is easier or harder to learn than the policy.
Multi-task RL treats which task you are on as part of the state, s = (s̄, z_i), so the problem is still an ordinary MDP and standard RL algorithms still apply. Lecture 12 of CS224R covers two kinds of sharing: weight sharing, where one network conditioned on z_i does every task, and data sharing via hindsight relabeling, where data collected for task A gets relabeled as data for task B. Goal-conditioned RL is the special case where the task is a goal state; relabeling with the state you actually reached eases the exploration problem of sparse rewards. Data sharing has three prerequisites: consistent dynamics across tasks, a reward you can evaluate, and an off-policy algorithm.
Meta-RL trains on many tasks so that a new task can be solved from a small amount of experience. Lecture 13 of CS224R frames it as "explore to collect a little data, then adapt using that data." The most direct approach is black-box meta-RL (RL²): a network with memory takes past (s, a, r) as input and keeps its hidden state across episodes. It is general and expressive but hard to optimize, especially when exploration is hard, because exploration and execution depend on each other and end-to-end training gets stuck. The slides then compare posterior sampling in PEARL, prediction-driven exploration in MetaCURE, and DREAM, which uses a task representation to train exploration and execution separately.
Long-horizon tasks are hard because the agent visits a huge number of states and has many chances to make mistakes or get stuck. Lecture 15 of CS224R answers with two levels: a high-level policy proposes subgoals, and a low-level policy runs at a higher frequency to reach them. The real design decisions are three: how to represent the subgoal, how to supervise each level, and when to switch to the next subgoal. The slides also admit that nobody has yet shown whether hierarchy beats a single policy with chain of thought.
Simulators are cheap, fast, and safe, and they hand you labels the real world never will, but they never match reality exactly. In Lecture 16 of CS224R, CMU's Guanya Shi sorts the ways to close that gap into three families: domain randomization trains one policy that works across many physical parameters; teacher-student trains a teacher on privileged information and then has a student that sees only real sensors imitate it; real2sim2real uses real data to make the simulator more faithful. The advanced topics are defining tasks from human motion data and choosing RL algorithms suited to sim2real.
VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.
The last CS224R lecture has three parts. It first folds the whole quarter into one toolbox. It then lists seven unsolved problems: domains without verifiable rewards, using prior data, world models, scaling, safety, hallucination and calibration, and evaluating generalist systems. Nearly half the deck is about how to do research: you need both an important problem and a workable plan, you front-load the risk, you consider pivoting early, and research only counts once you share it. Reading it alongside the 244 public 2026 final project reports shows what those principles look like in practice.