Skip to content
All tags

#cs224r

22 posts

CS224R L4: Actor-Critic and Value Estimation — Learn to Judge Good and Bad, Then Do More of the Good

Policy gradient can only judge good and bad from the rewards it actually received, so it wastes data. Actor-critic trains a second network, a value function (the critic), to estimate how good a state is, and uses it to compute advantages that weight the policy's (the actor's) gradient. There are three ways to estimate value: supervise directly with a rollout's summed rewards (Monte Carlo), supervise with this step's reward plus your own estimate of the next state (bootstrapping), or use an n-step return in between. L4 ends by pushing actor-critic off-policy, first by taking several gradient steps on one batch (where PPO starts) and then by reusing all past data from a replay buffer (where SAC starts).

Reading Stanford CS224R: A Guide to the Spring 2026 Deep Reinforcement Learning Course

CS224R is Chelsea Finn's deep reinforcement learning course at Stanford. It runs from imitation learning to RL for LLMs and robot foundation models. For Spring 2026, all 17 slide decks, the three homework handouts with starter code, and the default project spec with starter code can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the 2026 recordings are Canvas-only, the midterm and its solutions are not public, and HW2 and HW3 require Modal. The public recordings are from Spring 2025, so this series uses them as a supplement and flags the differences lecture by lecture.

CS224R Default Project: Fine-Tuning an LLM on Countdown with SFT, IPO, and RLOO

The CS224R Spring 2026 default project has you implement three stages on Qwen2.5-0.5B Base for the Countdown arithmetic reasoning task: SFT warm-start, IPO preference optimization, and RLOO with a rule-based verifier reward. All three are compared with the same vLLM evaluation, followed by a research extension of your choice. For the implementation, high-level trainers like SFTTrainer are banned, and so is any AI tool assistance; only the extension is exempt. The extension is half the grade for this project, and it's graded on methodology and documentation, not score. The starter code and datasets are public. What outside readers lack is Modal credits and the autograder.

CS224R L18: Open Problems in Deep RL, and How to Do Research

The last CS224R lecture has three parts. It first folds the whole quarter into one toolbox. It then lists seven unsolved problems: domains without verifiable rewards, using prior data, world models, scaling, safety, hallucination and calibration, and evaluating generalist systems. Nearly half the deck is about how to do research: you need both an important problem and a workable plan, you front-load the risk, you consider pivoting early, and research only counts once you share it. Reading it alongside the 244 public 2026 final project reports shows what those principles look like in practice.

CS224R L15: Hierarchical RL and Imitation Learning

Long-horizon tasks are hard because the agent visits a huge number of states and has many chances to make mistakes or get stuck. Lecture 15 of CS224R answers with two levels: a high-level policy proposes subgoals, and a low-level policy runs at a higher frequency to reach them. The real design decisions are three: how to represent the subgoal, how to supervise each level, and when to switch to the next subgoal. The slides also admit that nobody has yet shown whether hierarchy beats a single policy with chain of thought.

CS224R HW1: Regression BC, Flow Matching, and DAgger on Flappy Bird

Homework 1 of CS224R Spring 2026 tests imitation learning on a custom Flappy Bird environment. The policy predicts 20 future target heights at once and executes only the first 10. You implement MSE-regression behavior cloning, a flow matching policy, and DAgger, then compare them in easy and hard modes. The PDF, LaTeX template, and starter code are all public, and a CPU is enough to run it. Solutions, the autograder, and Gradescope are not public. This guide covers what each problem asks you to build and answer. It does not give solutions.

CS224R HW2: Gridworld Q-learning, PPO, and the Sawyer Hammer Task

CS224R Spring 2026 HW2 has three parts. First, tabular Q-learning on a 5×4 gridworld shows how reward design changes the learned path. Second, GAE plus PPO clipping tackles a hammer task that pays 1 only on completion. Third, an off-policy actor-critic with BC pretraining, a critic ensemble, and a higher UTD ratio, followed by a comparison of the two learning curves. The handout, starter code, and compute guide are public, but the assignment supports only Modal, and course credits go only to enrolled students.

CS224R HW3: AWAC, IQL, and Stitching on AntMaze

HW3 in CS224R (Spring 2026) has you fill in two offline RL algorithms, AWAC and IQL, and compare them on D4RL's AntMaze. Problem 1 runs AWAC on antmaze-umaze and antmaze-medium-diverse. Problem 2 compares IQL expectiles ζ = 0.2 and 0.9, runs the better value on medium-diverse, and then tests whether IQL can stitch a better path out of a PointMass dataset whose best return is only −46, against a filtered BC baseline that keeps the top 10% of trajectories. The PDF, LaTeX template, and starter code are public, but the assignment is meant to run on Modal, and course credits go only to enrolled students. This post covers the tasks and setup only, with no solutions.

CS224R L2: Imitation Learning and Policies That Can Represent Multimodal Distributions

Lecture 2 of CS224R Spring 2026 tackles two ways imitation learning fails. First, when demonstrations contain several reasonable behaviors, regression learns only their average. The fix is to make the policy a generative model (Gaussian mixtures, discretization plus autoregression, diffusion or flow matching) and to add action chunking. Second, compounding errors: once the policy slips, it reaches states the demonstrations never covered. The fix is DAgger or human-gated DAgger to collect corrections. The first two parts are exactly what HW1 covers.

CS224R L1: Framing Decision-Making as an RL Problem

The first lecture of CS224R Spring 2026 does three things: covers logistics, explains why deep RL is worth learning, and turns 'behavior' into something you can learn. The core is a set of definitions (state, action, trajectory, reward, policy) and one objective: maximize expected total reward. It ends on an example: fit ℓ2 regression to drivers where some change lanes and some go straight, and the policy learns their average, a half lane change nobody demonstrated. That problem is where L2 starts.

CS224R L13: Meta-RL, Teaching an Agent to Learn New Tasks Fast

Meta-RL trains on many tasks so that a new task can be solved from a small amount of experience. Lecture 13 of CS224R frames it as "explore to collect a little data, then adapt using that data." The most direct approach is black-box meta-RL (RL²): a network with memory takes past (s, a, r) as input and keeps its hidden state across episodes. It is general and expressive but hard to optimize, especially when exploration is hard, because exploration and execution depend on each other and end-to-end training gets stuck. The slides then compare posterior sampling in PEARL, prediction-driven exploration in MetaCURE, and DREAM, which uses a task representation to train exploration and execution separately.

CS224R L11: Model-Based RL, or Learn a Simulator and Don't Trust It Too Much

Model-based RL first learns a dynamics model that predicts s_{t+1}, then uses it in one of two ways: to generate extra training data (Dyna, MBPO) or to think a few steps ahead before acting (planning). The thread running through Lecture 11 of CS224R is how to avoid being dragged down by model error. Start synthetic rollouts from real states and keep them short, average errors out with an ensemble of models, and attach a value function to the tail of long-horizon plans. Whether a model is worth learning depends on whether it is easier or harder to learn than the policy.

CS224R L12: Multi-Task and Goal-Conditioned RL, Sharing Weights and Sharing Data

Multi-task RL treats which task you are on as part of the state, s = (s̄, z_i), so the problem is still an ordinary MDP and standard RL algorithms still apply. Lecture 12 of CS224R covers two kinds of sharing: weight sharing, where one network conditioned on z_i does every task, and data sharing via hindsight relabeling, where data collected for task A gets relabeled as data for task B. Goal-conditioned RL is the special case where the task is a goal state; relabeling with the state you actually reached eases the exploration problem of sparse rewards. Data sharing has three prerequisites: consistent dynamics across tasks, a reward you can evaluate, and an off-policy algorithm.

CS224R L5: Off-Policy Actor-Critic — the Shared Skeleton of PPO and SAC

PPO and SAC answer the same question: can you use an expensive batch of data more than once? PPO takes several gradient steps on one fresh batch and clips the new-to-old policy ratio to 1±ε. SAC keeps every past transition in a replay buffer and learns Q(s, a), so old data can still evaluate the current policy. PPO is stable and easy to tune. SAC is data-efficient and harder to tune.

CS224R Lecture 7: Offline RL, or Why Q-Learning Breaks When You Can't Collect More Data

Lecture 7 of CS224R (Spring 2026) asks how to learn a policy better than your data when all you have is a fixed dataset someone else collected. Running an off-policy algorithm like SAC on that data fails: the Q-function makes up values for actions the data never contains, and the policy goes looking for exactly those overestimated actions. The slides give two families of fixes. One trains the policy only on actions in the data (filtered BC, AWR, AWAC). The other uses an asymmetric expectile loss to estimate the value of a policy better than the data without ever querying out-of-data actions (IQL). Both can do something imitation learning can't: stitch good pieces of different trajectories together.

CS224R L3: Policy Gradients — Differentiating the Policy Without Knowing How the World Works

Policy gradient is the first online RL algorithm in CS224R. Its gradient looks almost exactly like the imitation learning gradient, except that each trajectory is weighted by its reward. Actions from good outcomes become more likely, and actions from bad outcomes become less likely. The raw version is very noisy, so L3 cuts the variance in two ways: count only future rewards (causality) and subtract the average reward (a baseline). It is also on-policy, so every gradient step needs fresh data. Importance sampling plus a KL constraint lets you take several steps on one batch.

CS224R L6: Q-learning and How to Stabilize It

Q-learning drops the actor from actor-critic: learn the optimal Q-function directly and act by taking the argmax. The price is that convergence is not guaranteed; even linear Q can diverge. Lecture 6 of CS224R pulls it back with three engineering tricks: a target network that holds the targets still, Double Q that separates choosing an action from valuing it to curb overestimation, and n-step returns that trade a little bias for speed.

CS224R Lecture 8: Where Rewards Come From, Learned from Examples and Preferences

Lecture 8 of CS224R (Spring 2026) spends a few slides wrapping up offline RL, then asks the question the first seven lectures skipped: where does the reward come from? Games have scores. Real robots, dialogue, and driving usually don't. The slides offer two routes. The first trains a goal classifier on success examples and uses it as the reward, but RL learns to exploit the classifier's blind spots; the fix is to keep adding states the policy visits as negatives, the same structure as a GAN. The second asks people which of two trajectories is better and learns a reward with the Bradley-Terry-style objective log σ(r(τw) − r(τl)), the same method LLM RLHF uses. The lecture's number-one takeaway is one line: rewards can't be taken for granted.

CS224R L17: RL for Robot Foundation Models (VLAs)

VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.

CS224R L10: RL for LLM Reasoning and Test-Time Compute

Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.

CS224R L9: RLHF, DPO, and Preference Optimization

Lecture 9 of CS224R Spring 2026 is a guest lecture by Archit Sharma, with slides adapted from CS224N. The spine is one chain of reasoning. Instruction tuning can't handle tasks with no right answer or errors of unequal weight, so we optimize human preferences directly. Human ratings are expensive and noisy, so we collect pairwise comparisons and fit a Bradley-Terry reward model. RLHF uses that model as the reward and runs policy gradient with a KL penalty. DPO uses the closed-form solution of the KL-constrained problem to write the reward as a log-ratio of policies, which turns the whole thing into a binary classification loss. The last part covers the frontier: reward hacking, verifiable rewards, and AI feedback in place of human feedback.

CS224R L16: Sim-to-Real Robot Learning

Simulators are cheap, fast, and safe, and they hand you labels the real world never will, but they never match reality exactly. In Lecture 16 of CS224R, CMU's Guanya Shi sorts the ways to close that gap into three families: domain randomization trains one policy that works across many physical parameters; teacher-student trains a teacher on privileged information and then has a student that sees only real sensors imitate it; real2sim2real uses real data to make the simulator more faithful. The advanced topics are defining tasks from human motion data and choosing RL algorithms suited to sim2real.