Skip to content

CS224R L12: Multi-Task and Goal-Conditioned RL, Sharing Weights and Sharing Data

Sep 30, 20261 min
TL;DRMulti-task RL treats which task you are on as part of the state, s = (s̄, z_i), so the problem is still an ordinary MDP and standard RL algorithms still apply. Lecture 12 of CS224R covers two kinds of sharing: weight sharing, where one network conditioned on z_i does every task, and data sharing via hindsight relabeling, where data collected for task A gets relabeled as data for task B. Goal-conditioned RL is the special case where the task is a goal state; relabeling with the state you actually reached eases the exploration problem of sparse rewards. Data sharing has three prerequisites: consistent dynamics across tasks, a reward you can evaluate, and an off-policy algorithm.

🌏 中文版

This post is based on the Spring 2026 edition of CS224R. It is part 16 of the Reading Stanford CS224R series. It follows L11 Model-Based RL and covers Lecture 12, "Multi-Task and Goal-Conditioned RL," on May 8, 2026. HW3 was due at 9 pm the same day.

Official sources used:

Access level is A3: the slides download anonymously, and the 2026 recordings live only on Canvas.

Companion video (supplementary): Spring 2025 Lecture 12: Multi-Task RL (about 70 minutes). The first half of the 2025 L12 slides is still wrapping up model-based RL (synthetic data generation and when to use model-based RL), so the start of the recording probably covers that too. This post follows the 2026 slides.

Why learn many tasks at once

Page 6 asks whether we can train a generalist policy that does many "tasks" rather than one. The examples cut across fields: an LLM assistant that books travel and buys groceries, a legged robot that walks, runs and dances, a mobile manipulator that hangs up a towel and unloads a dishwasher, a music recommender personalized to many users, and a game agent that plays Flappy Bird and Pokemon. These tasks may differ in reward, dynamics and even action space.

Page 7 recaps the course so far (imitation, on-policy, off-policy and offline RL, model-free and model-based RL, reward functions) and asks: what has been the biggest challenge? Data efficiency. The multi-task idea is to amortize the data cost across many tasks, because much of what gets learned can be shared: grammar for an LLM assistant, balance for legged robots, similarity between users for recommenders.

The slide also gives a deeper motivation: generalist ML systems are often more reliable and perform better than specialists.

The learning goal: how to share weights and data across tasks for learning efficiency.

Tasks are just different MDPs

Formalizing a task

Page 8 writes a task as an MDP: T_i ≜ {S_i, A_i, p_i(s_1), p_i(s'|s, a), r_i(s, a)}, meaning state space, action space, initial state distribution, dynamics and reward. The slide notes this definition allows far more variety than the everyday meaning of "task."

Pages 9–10 practice spotting which parts vary, using four examples:

ExampleWhat varies
Character animation learning multiple maneuversr_i
Recommending videos to multiple usersdynamics and r_i
Getting dressed with different garments and initial statesinitial state distribution and dynamics
Shirt folding across multiple robotsstate space, action space, initial state distribution, dynamics

How to tell the policy which task it is on

Page 11 lists three kinds of task identifier z_i: a task index (task 0, task 1, …), a language description, or a video showing what to do.

Pages 12–13 offer another view: make the task identifier part of the state, s = (s̄, z_i). Then it is still a standard MDP, so why not apply standard RL algorithms? The slide's answer: you can. In some cases you can do better.

The reward in multi-task RL is the same as before. In goal-conditioned RL, z_i is the goal state s_g and the reward is r(s) = −d(s̄, s_g), where the distance d can be Euclidean or a sparse 0/1.

Weight sharing: condition on z_i

Start with multi-task imitation learning

Page 14 borrows a trick from multi-task supervised learning: stratified sampling. Put data from every task into each minibatch and the gradients have lower variance.

Pages 15–18 show how robot policies take in z_i in practice:

Page 18 leaves a question open: what about high-level, long-horizon tasks like "clean the kitchen"? That waits for the lecture on hierarchy.

The basics of multi-task RL

Page 19 is short:

  • Policy: π_θ(a | s̄) becomes π_θ(a | s̄, z_i)
  • Q-function: Q_φ(s̄, a) becomes Q_φ(s̄, a, z_i)
  • Consider a replay buffer per task for stratified sampling

Then the slide asks the lecture's central question: if you collect data while conditioning on z_1, can you reuse it to learn task 2?

Data sharing: hindsight relabeling

The accidental good pass

Page 21 uses ice hockey. Task 1 is passing and task 2 is shooting goals. What if you try to shoot and accidentally make a good pass? Store the experience as normal, and also relabel it with the other task's ID and reward and store that copy. This is hindsight relabeling, also called hindsight experience replay (HER). The slide literally says "relabel experience with task 2 ID"; by the logic of the example, the accidental outcome is a pass, so the relabel target should be task 1 (passing). I read it by the logic and quote the original wording for comparison.

The multi-task algorithm

Page 22:

  1. Collect data D_k = {(s_{1:T}, a_{1:T}, z_i, r_{1:T})} with some policy
  2. Store it in the replay buffer
  3. Hindsight relabeling: relabel D_k for task T_j to get D'_k, with r'_t = r_j(s_t), and store it too
  4. Update the policy using the replay buffer

Which task j should you relabel for? The slide offers two options: pick randomly, or pick tasks in which the trajectory gets high reward.

The same page asks: in what scenarios can we apply relabeling? The slide's answer is three conditions:

  • The form of the reward function is known and can be evaluated
  • Dynamics are consistent across goals or tasks
  • You use an off-policy algorithm
Why it has to be off-policy (my note)

The slide lists the conditions without explaining them one by one. What follows is my inference from earlier lectures. Relabeled data was collected under the policy for task i, so for the task-j policy π(a | s̄, z_j) it is data from a different policy. On-policy methods like those in L3 policy gradients require data from the current policy. Off-policy methods like L6 Q-learning only need (s, a, r, s') transitions, so they can use this data directly. Consistent dynamics matter for the same reason: a transition (s, a, s') must still be one that can actually happen under task j.

The goal-conditioned algorithm

Page 23 applies the same recipe to goal-conditioned RL:

  1. Collect data D_k = {(s_{1:T}, a_{1:T}, s_g, r_{1:T})}
  2. Store it in the buffer
  3. Relabel using the last state reached as the goal: D'k = {(s{1:T}, a_{1:T}, s_T, r'_{1:T})}, with r'_t = −d(s_t, s_T), and store it too
  4. Update the policy

Other relabeling strategies? The slide's answer: use any future state from the trajectory. The result: exploration challenges are alleviated. This part cites both Kaelbling's Learning to Achieve Goals (IJCAI 1993) and the HER paper.

Try this: draw a one-dimensional corridor of five cells on paper, with the goal at the right end. The reward is 0 on reaching the goal and −1 otherwise. Write down a trajectory that only makes it to cell 3, then label its rewards twice: once with the original goal and once with the last state as the goal. With the original labels every step is −1. After relabeling, at least one step is 0, and that is the moment the Q-function first gets a useful signal.

Tying it together: what each kind of sharing needs

Pages 26–28 summarize:

  • Multi-task RL is single-task RL in a joint MDP whose state is s = (s̄, z_i); each episode first samples a task
  • Goal-conditioned RL is a special case with z_i = s_g, where every task means reaching some goal state. The reward is δ(s = s_g) for discrete states and δ(‖s − s_g‖ ≤ ε) for continuous states

Pros and cons of goal-conditioned RL:

  • No need to define a reward (self-supervised)
  • Many tasks can be framed as goal reaching
  • Can be fairly hard to train
Weight sharingData sharing
HowTrain one network to do all tasks, conditioned on z_iAdd data collected for one task to another task's buffer by relabeling the reward and task identifier
Requires—Same dynamics across tasks, evaluatable rewards, an off-policy algorithm
Goal-conditioned—Directly applicable

The next lecture asks whether we can adapt quickly to a new task, in other words in-context learning for RL. See L13 Meta-RL.

Follow-up exercise: Spring 2025 HW4 Part 1 (archived)

2026 had only three assignments, and none covers this lecture. Part 1 of the 2025 HW4 matches the topic exactly. Its PDF and starter code still download anonymously, which makes it good practice for self-learners. Below are the requirements only, with no solutions.

Per the PDF, Part 1 has four steps:

  1. Adapt an existing DQN to be goal-conditioned, with a Q-network that takes the concatenated state and goal
  2. Run goal-conditioned DQN on two environments
  3. Implement HER on top of it
  4. Compare performance with and without HER

The two environments:

  • Bit flipping: the state and goal are binary vectors of length n, and each step flips one bit. The reward is −1 when state and goal differ and 0 when they match, a sparse-reward example; the larger n is, the rarer non-negative rewards become
  • 2D Sawyer reach: move a robot arm's end effector to a goal XY position, with reward equal to negative Euclidean distance, a dense-reward example

The functions to implement live in run_episode.py and trainer.py. HER comes in three variants to implement: final, random and future (future may only pick goals from states after the current step). The analysis questions scale the bit count from 6 to 15 to 25 and compare runs with and without HER, then compare the three variants at 15 bits, and finally compare how much HER contributes in bit flipping versus Sawyer reach.

Things self-learners should watch for:

  • The assignment requires AWS EC2. The PDF specifies a c4.4xlarge instance with a course-provided custom AMI and says other platforms are not supported. Readers outside the course have to set up an environment locally from the starter code's README
  • The PDF prohibits using generative models to help write code for this assignment
  • The autograder and Gradescope are not public, so you judge results from your own tensorboard curves

Further reading

What this post can and cannot confirm

Confirmed: the text, algorithms and summary table of the 2026 slides, the schedule date and assigned reading, the contents and compute rules of the 2025 HW4 PDF, and the title and length of the 2025 L12 video. Not confirmed: the in-class discussion of the question on page 9, the video material on pages 15–18, and how the three relabeling prerequisites were explained aloud (the collapsible section above is my inference).

Series navigation: previous L11 Model-Based RL | next L13 Meta-RL | Series overview

References