🌏 中文版
This post is based on the Spring 2026 edition of CS224R. It is part 17 of the Reading Stanford CS224R series. It follows L12 Multi-Task and Goal-Conditioned RL and covers Lecture 13, "Meta-RL," on May 13, 2026. The in-class midterm followed two days later, on May 15.
Official sources used:
- The 45-page slide deck 13_cs224r_metarl_2026.pdf
- The assigned reading on the schedule: RL²: Fast Reinforcement Learning via Slow Reinforcement Learning (Duan et al. 2016). The schedule row for the May 15 exam also lists Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices (Liu et al. 2021, i.e. DREAM)
- Follow-up exercise: Part 2 of Spring 2025 HW4 (archived from 2025)
Access level is A3: the slides download anonymously, and the 2026 recordings live only on Canvas.
Companion videos (supplementary):
- Spring 2025 Lecture 13: Meta RL (about 69 minutes), matching the problem setup and black-box methods in the first half of this lecture
- Optional: Spring 2025 Lecture 14: Exploration (about 73 minutes)
2026 dropped the 2025 L14 lecture on exploration; slot 14 became the exam day. Comparing the two years' slides: the first half of the 2025 L14 slides covers exploration algorithms for bandits and exploration in robotics and LLMs, and the second half covers learning to explore via meta-learning, which is where DREAM and the CS education application live. 2026 folded that second half into this lecture, while bandit exploration does not appear anywhere in the 2026 slides. If you want the spoken walkthrough of DREAM, the recording to watch is 2025 L14.
The setting: a new espresso machine
Page 7's example: you want a robot to learn a new espresso machine. Training it from scratch with PPO would take millions of attempts. A person with a little instruction, or with prior coffee experience, learns in minutes. People don't start from scratch: experience with other espresso machines, or just general motor control, helps. The same goes for math: past problem-solving experience helps you solve a hard new problem faster.
Page 8 sorts "use experience from earlier tasks to learn a new one" into three framings (the slide notes it is adapted from Sergey Levine):
| Framing | How | Requirement |
|---|---|---|
| Forward transfer | Train on a source task, then fine-tune on the target task | Source and target tasks must be similar |
| Multi-task transfer | Train on many tasks, transfer to a new one; the task descriptor z_i must capture task structure for zero-shot transfer | The target task must resemble the training task distribution |
| Meta-learning | Learn to learn on many tasks; training accounts for the fact that you'll adapt to a new task, and the task descriptor is a few examples given in context | Same as above |
The lecture's two learning goals: understand the meta-RL problem statement and setup, and understand the basics and challenges of black-box meta-RL algorithms.
Intuition: it's the same thing as in-context learning
Pages 9–10 step outside RL first. Show someone a few paintings by Braque and Cézanne, then ask who painted a new one: that's few-shot image classification. In-context learning in LLMs has the same shape: train a model ŷ = f(x, D_train) that predicts from a few examples.
Pages 11–12 switch to maze navigation in RL (diagram adapted from Duan et al. 2017):
- At meta-test time: in a new maze, first collect a little experience D_tr with an exploration policy π^exp, then use D_tr to get a policy π^task that solves this maze
- At meta-train time: learn how to explore efficiently and how to solve many mazes, which means learning π^exp and π^task together
The slides flag two points. The exploration-exploitation trade-off is a unique part of meta-RL. And the key assumption is that meta-test MDPs come from the same task distribution as meta-training MDPs, which is why we can expect generalization.
Page 13's task distributions: navigating different mazes, walking on different terrains and slopes, manipulating different objects toward different goals, and dialog with users who have different preferences.
Mechanism 1: the problem setup
Page 14 puts three problems side by side:
- Multi-task learning: solve several tasks at once, minimizing Σ L_i(θ, D_i)
- Transfer learning: after solving source task a, solve target task b using what was learned from a
- Meta-learning: given data from tasks T_1…T_n, quickly solve a new task T_test
An RL task is defined as in the previous lecture: T_i ≜ {S_i, A_i, p_i(s_1), p_i(s'|s, a), r_i(s, a)}.
Page 15 has two variants:
- Episodic: the input is k rollouts from the exploration policy; the output is π^task(· | s, D_train)
- Online: the input is the first 1…k timesteps from the exploration policy; the output is again π^task(· | s, D_train)
The slide notes that D_train can be viewed as a task identifier z_i, and that π^exp and π^task could share parameters.
Mechanism 2: black-box meta-RL
A memory network that takes reward as input
Page 17's architecture: a black-box network with memory (a transformer or a neural net with memory) takes (s_t, r_{t−1}) at each step and outputs a_t. D_train is this ever-growing history.
The slide asks: how is this different from simply doing RL with a recurrent policy? Three answers:
- The reward is passed as input
- It is trained across multiple MDPs
- The hidden state is maintained across episodes within a task
The algorithm
Pages 18–19:
Meta-training
- Sample a task T_i
- Roll out π(a | s, D_i^tr) for N episodes under T_i's dynamics and reward
- Store the sequence in T_i's replay buffer
- Update the policy to maximize discounted return across all tasks
Meta-test
- Sample a new task T_j
- Roll out π(a | s, D_j^tr) for up to N episodes
Note that meta-test involves no gradient updates. Adaptation happens entirely in the hidden state.
Page 20: three architectures and optimizers
| Paper | Architecture | Outer optimizer |
|---|---|---|
| RL² (Duan et al. 2017), Learning to Reinforcement Learn (Wang et al., CogSci 2017) | RNN | TRPO/A3C (the slide notes: similar to PPO) |
| SNAIL (Mishra et al., ICLR 2018) | Attention + 1D convolution | TRPO |
| PEARL (Rakelly et al., ICML 2019) | Feedforward + averaging | SAC (off-policy with a replay buffer) |
Three examples
- Example 1 (pages 21–22, SNAIL): learning to navigate a maze visually, training on 1000 small mazes and testing on held-out small mazes and on large mazes
- Example 2 (page 23, Qu et al., Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning): the tasks are different math problems; the model explores different strategies, then answers using the best one. The key idea is optimizing test-time compute. The slide reports "higher performance with same # of tokens" and "same accuracy with 1.6x fewer tokens." The slide labels the paper ICML 2019, but the arXiv version was posted in March 2025, so go by the arXiv date
- Example 3 (page 27, PEARL): continuous control where tasks differ in direction, velocity or physical dynamics. The slide says meta-RL algorithms are very efficient at new tasks, but what about meta-training efficiency? It asks whether off-policy meta-RL should be more or less efficient than on-policy meta-RL
Another view: it's really a POMDP
Page 24 connects meta-RL to multi-task policies. A multi-task policy is π_θ(a | s, z_i); meta-RL is a multi-task policy that uses experience as the task identifier. What about goal-conditioned policies? The slide's answer: rewards are a strict generalization of goals, and meta-RL aims to adapt to new tasks (k-shot) whereas goal-conditioned policies generalize to new goals (0-shot).
Pages 25–26: when z_i isn't in the state, the agent is unsure which MDP it is in and has to infer z_i from experience. So meta-RL can be seen as exploring to infer the unknown task, then performing it: a special kind of partially observed MDP (POMDP).
Page 28 sums up black-box meta-RL:
- General and expressive
- Many design choices in architecture
- Hard to optimize
- Inherits sample efficiency from the outer RL optimizer
Try this: take any RL codebase with a recurrent policy and list the three changes that turn it into RL²: concatenate the previous reward into the input, don't reset the hidden state at episode boundaries, and switch to a new task for each "trial." Once all three are in, it has gone from ordinary RL to meta-RL.
Mechanism 3: should exploration and execution be separated?
Why end-to-end is hard
Page 30 gives the first approach: optimize exploration and exploitation end-to-end with respect to task reward (RL² and a line of related methods). It is simple and in principle yields the optimal exploration-exploitation trade-off. The downside is a challenging optimization when exploration is hard.
Pages 31–32 illustrate with hallways. There are N hallways, different tasks require reaching the end of different hallways, and a sign nearby says where to go:
| Episode during meta-training | Training signal |
|---|---|
| Walk straight to the end of the correct hallway | Positive reward for the current task, but D_tr looks the same as for any other task |
| Go to the wrong hallway, then the right one | Signal about a suboptimal exploration-plus-exploitation strategy |
| Look at the sign | Good exploratory behavior, but this behavior earns no reward at all |
The slide's conclusion: it's hard to learn exploration and exploitation at the same time.
Pages 33–34 switch to kitchens: you learned cooking tasks in previous kitchens and need to pick up tasks quickly in a new one. End-to-end training hits a chicken-and-egg problem. If you can't find the ingredients (bad exploration), you can't learn to cook (bad execution); if you can't cook, every exploration strategy earns low reward. The slide calls this the coupling problem, which can lead to poor local optima and poor sample efficiency. This page cites the DREAM paper.
Three alternatives
Pages 36–39:
| Method | How | Upsides | Downsides |
|---|---|---|---|
| 2a. Posterior sampling (PEARL, also called Thompson sampling) | Learn a distribution over a latent task variable, p(z) and q(z | D_tr), plus task policies π(a | s, z); sample z from the current posterior, then act with π | Easy to optimize; many are based on principled strategies | The slide's counterexample: with goals far away and a sign on the wall naming the correct goal, posterior sampling never goes to read the sign |
| 2b. Predict dynamics and reward (MetaCURE, Zhang et al. 2020) | Train a model f(s', r | s, a, D_train) and collect D_train so the model becomes accurate | Same as above | Breaks down with many distractors or complex, high-dimensional state dynamics |
| 2c. Predict a compressed task representation (DREAM) | Train f(z_comp | D_train) and collect D_train so task prediction becomes accurate | Optimal strategy in principle, and easy to optimize in practice | Requires a task identifier during training |
The slide's overall verdict on 2a and 2b: in some environments they can be suboptimal by an arbitrarily large amount.
How DREAM actually trains (per the Spring 2025 HW4 Part 2 PDF)
The 2026 slides stop at the table cell above. The 2025 HW4 PDF spells it out in more detail; the summary below follows that PDF:
- Each meta-training task has a unique problem ID μ; at meta-test time DREAM does not assume access to it
- Learning execution: an encoder F_ψ(z | μ) turns μ into a task representation z, and the execution policy π^task(a | s, z) conditions on z to maximize return. The encoder is trained so that z holds only the information needed to solve the task
- Learning exploration: the exploration policy π^exp tries to maximize mutual information between the exploration trajectory τ^exp and z, optimized through a variational lower bound with a decoder q_ω(z | τ^exp)
- Splitting the bound into per-step differences gives the exploration policy an intrinsic reward: how much each newly observed (a_t, r_t, s_{t+1}) raises log q_ω(z | τ), i.e. the "information gain" about z from that step
- At meta-test time, the decoder turns the exploration trajectory into z and hands it to the execution policy
An application: grading student programs with meta-RL
Pages 40–43 are a "time permitting" section. Grading interactive student programs (Code.org's Bounce, CS106A's Breakout) eats TA time, so meta-RL can learn how to explore a program to find its bugs. The slides report that AI-assisted grading in CS106A (Spring 2023) was 44% faster and 6% more accurate, that Stanford TAs like using it, and that the autograder prepopulates the rubric and shows videos. The cited papers are Liu et al. at NeurIPS 2022 and at SIGCSE 2024.
Tying it together: from multi-task to meta-RL
Read the last three lectures as one arc:
- L11 learns a model of the environment
- L12 tells the policy what the task is (z_i) and shares weights and data across tasks
- L13 doesn't even tell the policy what the task is, and lets it infer the task from experience
Page 45 says the next lecture (the following Wednesday) is on hierarchy. See L15 Hierarchical RL and IL.
Follow-up exercise: Spring 2025 HW4 Part 2 (archived)
Per the PDF, Part 2 asks you to:
- Experiment with a black-box meta-RL method (RL²) trained end-to-end
- Implement components of DREAM to replace the end-to-end objective
- Compare end-to-end and decoupled optimization
The environment is a grid world with buses. Each episode the agent gets a goal and should reach it as fast as possible; riding a bus teleports it to the other bus of the same color. Tasks differ only in how the corner buses are permuted, for 4! = 24 tasks, and a map at a fixed location reveals each bus's destination when the agent stands on it. The setting has one exploration episode and one exploitation episode, and only the exploitation return counts.
Structure of the problems:
- Problem 0: pen-and-paper analysis of the return from walking directly, the best return when bus destinations are known, the exploration policy that finds all destinations in the fewest steps, and the best return a meta-RL agent can achieve
- Problem 1: run RL² (the PDF estimates 2–3 hours), read tensorboard and the videos, and describe the exploration and exploitation behavior it learns and whether it reaches the optimum
- Problem 2: implement the decoder loss and the exploration intrinsic reward in
encoder_decoder.py, run DREAM, and compare it with RL²
Things self-learners should watch for: as with Part 1, the PDF requires the course's AWS EC2 setup, doesn't support other platforms, and prohibits using generative models to help write the code. The starter code still downloads anonymously, but you set up the environment yourself and there is no autograder.
Further reading
- Berkeley CS285 Spring 2026: harder exploration and reuse across tasks, for how another course covers exploration and meta-learning
- CS224R L10: RL for LLM reasoning, which shares the test-time compute theme of this lecture's Example 2
What this post can and cannot confirm
Confirmed: the text, tables and citations of the 2026 slides, the schedule date and assigned readings, how the 2025 L13 and L14 slides split the material, the contents of the 2025 HW4 PDF, and the titles and lengths of the two 2025 videos. Not confirmed: the in-class answers to the questions on pages 17 and 27, the full experimental setup behind Example 2 (the slide has two lines of results), and whether the 2026 lecture covered bandit exploration aloud.
Series navigation: previous L12 Multi-Task and Goal-Conditioned RL | next L15 Hierarchical RL and IL | Series overview
References
- CS224R: Deep Reinforcement Learning (Spring 2026 course site and schedule)
- Lecture 13 slides: Meta Reinforcement Learning (2026)
- Spring 2025 Lecture 13: Meta RL (YouTube, supplementary)
- Spring 2025 Lecture 14: Exploration (YouTube, optional)
- Spring 2025 Lecture 14 slides: Exploration (archived)
- CS224R Spring 2025 Homework 4 (archived)
- CS224R Spring 2025 HW4 starter code (archived)
- Duan et al. RL²: Fast Reinforcement Learning via Slow Reinforcement Learning (arXiv 1611.02779)
- Liu, Raghunathan, Liang, Finn. Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices (arXiv 2008.02790, ICML 2021)
- Rakelly, Zhou, Quillen, Finn, Levine. Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables (arXiv 1903.08254, ICML 2019)
- Mishra, Rohaninejad, Chen, Abbeel. A Simple Neural Attentive Meta-Learner (arXiv 1707.03141)
- Wang et al. Learning to Reinforcement Learn (arXiv 1611.05763)
- Qu et al. Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning (arXiv 2503.07572)
Loading...