🌏 中文版
This post is based on the Spring 2026 edition of CS224R. It is part 20 of the Reading Stanford CS224R series. It follows L16 Sim-to-Real Robot Learning and covers Lecture 17, "RL for Robots: RL for VLAs," on May 27, 2026 (Wednesday of week 9). The slide deck's cover title is "RL for Robot Foundation Models."
Official materials used:
- The 2026 slides, 17_cs224r_rl_vlas_2026.pdf (34 pages)
- The schedule lists no reading for this lecture; every paper below is one the slides cite
Access level is A3: the slides download anonymously, and the 2026 recordings live on Canvas.
There is no companion video. This lecture is new in 2026, with no counterpart on the Spring 2025 archive or the 2025 playlist. This post relies on the slides alone and does not fill in what the figures and videos leave unsaid.
Slide 4 also sets expectations: this is an open, active research problem, and the lecture covers some recent themes plus the speaker's opinion on the area. Read it as a research map, not settled knowledge.
Setting: from simulation to pretrained models
Slide 3 connects this lecture to the last:
- Last lecture: can we do RL on simulated robots and transfer behaviors to the real world?
- Today: how do we do RL on real robots with pretrained foundation models?
What a VLA looks like
Slides 5–6 describe the most common robot foundation model, the vision-language-action (VLA) model:
- It starts from a pretrained vision-language model (VLM); an alternative design starts from a generative video model.
- The training mixture often includes robot demonstrations (imitation learning), VLM tasks (question answering, captioning, detection), and human video (motion prediction).
- It often includes a diffusion-based action expert that:
- predicts continuous actions with diffusion or flow matching
- attends to all activations of the LLM backbone
- is designed to avoid multiple forward passes through the whole backbone
- often does not pass gradients to the backbone
You implemented flow matching in HW1. For more intuition, see the MIT 6.S184 flow matching and diffusion guide.
The problem: stuck at 80%
Slide 7: VLAs trained with imitation learning often plateau around 80%. The figure shows π0.5 in unseen rooms.
- This mirrors how LLMs get better with RL after SFT (see L9 RLHF and Preference Optimization).
- Robots acting autonomously often need 99%+ reliability.
- So this is a natural use case for RL fine-tuning, and a pretrained VLA can serve as an effective initialization.
The slide also notes that DAgger helps and is often used alongside RL.
Why RL for VLAs is hard
Slide 8 gives two reasons.
1. VLAs are large, so gradient updates are expensive
- You often want to train in the cloud while robot rollouts run on a local computer.
- Each experiment takes longer to iterate.
- Extensive hyperparameter tuning is expensive and slow.
2. VLAs are pretrained with imitation learning
- There is no pretrained value function or critic.
- They are often trained with diffusion or flow matching, which makes off-the-shelf RL algorithms harder to apply.
- They are often trained with action chunking, also an uncommon choice for RL.
Online or iterated offline
Slide 9 compares the two loops:
| Online RL (e.g. SAC) | (Iterated) offline RL | |
|---|---|---|
| Data per round | 1 timestep | 1k episodes |
| Gradient steps per round | 1 | 10k |
| When there's a bug or a bad hyperparameter | Rerun the experiment and recollect data | Rerun training on the existing dataset |
The slide's conclusion: offline is much simpler for large models. Learning rate, number of epochs, gradient clipping: when you get these wrong offline, you just retrain. Offline RL basics are in L7 Offline RL.
Try this: if you're deciding whether to RL-fine-tune a robot or agent model, first estimate how much time and labor one round of data collection costs. The bigger that number, the more you should start with iterated offline RL.
Tool 1: can we just use PPO?
Slide 11 answers "yes, but." Two examples:
- Fine-tuning OpenVLA with RL: Li et al., SimpleVLA-RL (2025)
- Fine-tuning π0.5 with RL: Chen et al., πRL (2026)
The two limits:
- It requires a massive number of online policy rollouts, and many papers don't even report the sample count.
- Results are limited to simulation-based training.
Tool 2: iterated offline RL as supervised learning
Slide 12 introduces key theme #1: can we build a method on supervised learning? If so, it may scale more easily to large models and datasets.
Two parts:
- Learn a value function: fit V with Monte Carlo.
- Use it to get a better policy: supervise the policy to take the actions V thinks are better.
The slides use Physical Intelligence's π*0.6 (2025):
- Value function (slide 13): fit a multi-task, language-conditioned V on a large demonstration dataset. It uses a pretrained VLM, is conditioned on current images, the language prompt, and episode metadata, and predicts time to go.
- Policy (slide 14): advantage-conditioned supervised learning.
- Estimate the advantage A(s, a) from predicted values.
- Binarize the advantage to tell the policy whether the action was good or bad.
- Fine-tune the policy with supervised learning, conditioned on the binarized advantage.
- Full algorithm (slide 15): collect a large batch of rollouts and interventions → update the value function to predict time-to-go → update the VLA with advantage-conditioned supervised learning, and repeat.
Slide 16's headline: RL post-training gives 2x the throughput of IL post-training. Slide 17's videos show a robot making a latte with a person and making lattes reliably through 13 hours of operation.
Why conditioning on a binarized advantage improves the policy
This is my addition; the slides don't spell it out. During training the policy sees actions tagged "good" and "bad" and learns the action distribution under each tag. At inference you always pass "good," so it samples only from the good-action distribution. The whole loop stays supervised, with no policy gradient, which sidesteps slide 8's problem that diffusion and flow matching resist standard RL. The idea is close to the advantage-weighted methods from L7: both use the advantage to filter or weight the supervision signal.
Is this the best recipe?
Slide 18 lists the speaker's own reservations:
- TD updates should be able to beat Monte Carlo value learning, even at large scale.
- The method should also benefit from more powerful policy improvement.
- Online RL should be more data efficient and reach higher performance by seeking out failure modes and ruling out new strategies faster, at the cost of more infrastructure.
Tool 3: online RL by reducing dimensionality
Slide 20 introduces key theme #2: instead of fine-tuning the VLA end to end, learn a separate Gaussian policy using the VLA's representation. Two versions:
- Version A: treat the VLA's diffusion noise as an action space and train an RL policy to control it (Wagenmaker et al., DSRL, 2025).
- Version B: compress the VLA's visual representation and do RL on top of it (Xu et al., RLT, 2026).
A side note on the slide: after RL, you can distill the policy's data back into the VLA.
DSRL: a policy that picks the noise
Slides 21–23 expand version A, from Steering Your Diffusion Policy with Latent Space Reinforcement Learning (Wagenmaker, Nakamoto, Zhang et al., CoRL 2025). The intuition: different noise vectors denoise into different actions, so train a policy to output noise that leads to good actions.
Sampling:
- Sample w_t ∼ π_steer(· | s_t; θ).
- Denoise an action chunk a_{t:t+h} = π_VLA(s_t, w_t).
- Run a_{t:t+h} in the environment and observe s_{t+h}.
Training:
- Collect a rollout (s_1, w_1, r_1, …, s_T) and add it to the buffer.
- Sample a minibatch of transitions from the buffer.
- Update Q_ϕ(s_t, w_t) and π_steer(w_t | s_t; θ) with SAC.
Slide 23's numbers: 65 online episodes, about 10k steps, and roughly O(100x) more efficient than PPO.
Note that Q takes the noise w as input, not the actual action a. The VLA is frozen and treated as part of the environment, so SAC only has to handle a low-dimensional Gaussian policy. That sidesteps both of slide 8's problems at once: the model's size and diffusion's resistance to RL. SAC is covered in L5 Off-Policy Actor-Critic.
Tool 4: online RL with a small policy that edits actions
Slide 25 introduces key theme #3: learn a small Gaussian policy that edits the (diffusion) VLA's actions. Two versions:
- Version A: an actor-critic algorithm, then distill back into the VLA (Xiao et al., Probe-Learn-Distill, 2025)
- Version B: an actor-critic algorithm plus best-of-N sampling at test time (Dong et al., EXPO-FT, 2026)
EXPO: an edit policy plus sampling
Slides 26–30 walk through EXPO: Stable Reinforcement Learning with Expressive Policies (Dong, Li, Sadigh, Finn, ICLR 2026). The basic recipe:
- Optimize a smaller Gaussian edit policy to maximize Q-values, like typical RL.
- Train the base policy with imitation on all successes.
The slide warns that on its own this may not be stable: the edit policy naturally lags behind the Q-function, and it could collapse.
The stabilizer is to maximize the latest Q-function on the fly via sampling (slide 27):
- Sample multiple a from π_base.
- Sample multiple ã from π_edit(ã | s, a).
- Pick the one in {a_1, …, a_n, ã_1, …, ã_n} with the highest Q.
This reduces lag behind the Q-function and is resilient to edit-policy collapse. The speaker asks whether it can be viewed as a form of test-time scaling.
Slide 28 carries a "!!" note: when fitting Q, how do you pick a′ in the Bellman backup? Use the on-the-fly policy, the same sample-then-pick-highest-Q procedure.
Slide 30's ablations:
- No edit policy: value maximization is significantly hindered.
- No on-the-fly policy in the Bellman backup: poor performance in some environments.
Slide 29 shows real-robot results from EXPO-FT (Dong, Hung, Gao, Sadigh, Finn, 2026, Sample-Efficient RL Fine-Tuning for VLAs):
- Higher reliability than SFT and DAgger
- Trained on 19 minutes of experience on average, about 11k steps
- Learns more efficiently and effectively than DSRL and HIL-SERL
Try this: if all you have is a frozen VLA, or any frozen generative policy, this lecture gives you two entry points that leave its weights alone: control its input noise, or add a small corrector on its output. Start with whichever interface your stack exposes most easily.
Tying it together: summary and outlook
Slide 32's summary:
- Challenges: VLAs are large, so gradient updates are expensive; VLAs are pretrained with imitation learning.
- Three themes:
- #1: build on supervised learning for scalability (offline RL)
- #2: learn a separate Gaussian policy on the VLA's representation (online RL)
- #3: learn a small Gaussian policy that edits the (diffusion) VLA's actions (online RL)
Slide 33 is the speaker's outlook, in two parts:
- Exciting progress: evidence that RL can substantially improve the performance and speed of state-of-the-art VLAs, and evidence of reaching the performance needed for real-world deployment.
- No satisfying solution yet: online RL should be more efficient and effective than offline, and needing residual policies or dropping to a latent space seems unsatisfying compared with doing RL directly on the VLA's weights.
This closes the loop on slide 41 of L15 Hierarchical RL and Imitation Learning, which called RL fine-tuning of large (hierarchical) robot systems an open and important research direction.
The final lecture covers a summary, open problems, and how to do RL research. See L18 Frontiers and How to Do Research.
Further reading on this site:
- Berkeley CS285: inference and offline RL
- CS336 RLVR, the LLM-side counterpart of "RL after SFT"
What this post can and cannot confirm
Confirmed: the text, algorithm steps, numbers (80%, 99%, 2x, 13 hours, 65 episodes / ~10k steps, O(100x), 19 minutes / ~11k steps), and paper labels in the 2026 slides; the schedule's date and the absence of a reading; and that neither the 2025 archive nor the playlist has a matching lecture. Not confirmed: who gave this lecture (the slides name no one and just say "my opinion"); the specific values, baselines, and setups behind each figure; and the full papers for πRL, RLT, Probe-Learn-Distill, and EXPO-FT, which I did not find or open and cite only from the slide labels.
Series navigation: previous L16 Sim-to-Real Robot Learning | next L18 Frontiers and How to Do Research | Series overview
References
- CS224R: Deep Reinforcement Learning (Spring 2026 course site and schedule)
- Lecture 17 slides: RL for Robot Foundation Models (2026)
- CS224R Spring 2025 archive (no matching lecture)
- Physical Intelligence 2025: π*0.6: a VLA That Learns From Experience
- Wagenmaker et al. 2025: Steering Your Diffusion Policy with Latent Space Reinforcement Learning (DSRL)
- Dong et al. 2025: EXPO: Stable Reinforcement Learning with Expressive Policies
- Li et al. 2025: SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Luo et al. 2024: Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)
- Physical Intelligence 2025: π0.5
Loading...