Skip to content

CMU 11-868 RLHF Systems and Assignment 7: A VERL-Style Pipeline with a Reward Model, GAE, and PPO

Sep 30, 20261 min
TL;DR11-868's RL systems lecture has no slides; the Syllabus lists just one paper, ReaLHF. Assignment 7, on the other hand, is fully public. You train a DistilBERT reward model on Anthropic's HH-RLHF data (40 points), fill in GAE, the PPO loss, and entropy in a VERL-style trainer to fine-tune GPT-2 (40 points), and compare reward distributions before and after RLHF (20 points). The starter trainer never imports the verl package. What you learn is the RLHF dataflow, not VERL's distributed engine.

🌏 中文版

Version note: This post is based on the Spring 2026 offering of CMU 11-868 LLM Systems. Lecture details come from the Syllabus. The assignment page lives on the homework site, shared across terms, and the starter code is in llmsys_hw7, both as seen on 2026-09-30. The repo's last commit is 2026-05-02, and Fall 2026 hasn't changed it yet. Access grade: the assignment is A3, with the problems, starter code, tests, and grading criteria all public. The RL lecture itself is only A1: no slides, no video, just one paper.

Series: previous L26–L30 serving at scale: prefill/decode disaggregation, KV cache, and heterogeneous hardware | this is the last post in the series | Series overview

The course description lists "efficient implementation of RLHF," and this lecture plus Assignment 7 is where it lands. The two are public to very different degrees. The lecture has a title and one reading. The assignment has a full write-up, starter code, and grading criteria. This post covers what you can read for the lecture, then what the assignment asks of you.

No solutions here. I also don't derive the RLHF algorithms themselves (PPO, reward model theory); links to other courses are at the end.

The 4/15 lecture: a title and one paper

The 4/15 entry on the Syllabus is "Efficient Reinforcement Learning System for LLMs." That row has no [slides] link, and deck number 25 is missing from the file sequence, between the 4/13 vLLM deck (L24) and the 4/20 Dynamo deck (L26). The Fall 2026 Syllabus puts it on 11/23, again without slides.

The only material is the reading, the ReaLHF paper by Mei et al. (the paper calls the system ReaL). So what follows sticks to what the abstract supports.

The problem: supervised training has one model doing one kind of work. One round of RLHF involves several LLM instances. The actor, critic, reward, and reference models each generate, run inference, or train, and they depend on one another. Reusing the fixed parallelism of supervised training for all of this is inefficient.

The approach: redistribute parameters during training so each workload gets its own parallelism. The paper calls this parameter reallocation. A search algorithm with a lightweight runtime estimator finds an "execution plan" that says which GPUs each workload uses and how it's parallelized. The runtime engine then moves parameters according to the plan.

The result: the abstract reports up to 3.58x speedup over baselines on LLaMA models up to 70B parameters and 128 GPUs.

It's worth reading this against the previous post. Prefill and decode were split because they behave differently. Generation and training inside RLHF also behave differently, and ReaLHF's answer is to switch the same parameters between two parallel layouts.

The assigned papers: HybridFlow and VERL

The first two essential readings on the Assignment 7 page are the VERL docs and the HybridFlow paper. VERL is HybridFlow's open-source implementation; the assignment page expands the name as "Volcano Engine Reinforcement Learning."

The HybridFlow abstract makes its case in three steps:

  1. Classic RL can be drawn as a dataflow: nodes are neural-network computations, edges are data dependencies. RLHF turns every node into a distributed LLM training or generation program and every edge into many-to-many data transfer.
  2. A single controller that directs all computation and communication has too much dispatch overhead. Existing RLHF systems use multiple controllers instead, but that nests computation inside communication and makes algorithms awkward to change.
  3. HybridFlow mixes the two. Hierarchical APIs wrap computation and data dependencies so algorithms are easy to write and device mappings easy to change. A 3D-HybridEngine reshards the actor model between the training and generation phases with no redundant memory.

The abstract reports 1.53–20.57x higher throughput. The VERL docs currently list algorithms such as PPO, GRPO, and DAPO, and rollout backends such as SGLang and TensorRT-LLM. One page covers offloading rollout KV cache through Mooncake Store, which ties straight back to Mooncake in the previous post.

What Assignment 7 asks you to do

The assignment is titled "Introduction to RLHF." The task is to build a "VERL-like" framework and use RLHF to fine-tune a small model toward more helpful, more harmless answers. The data is Anthropic/hh-rlhf on Hugging Face, and the example data-prep command takes 10,000 samples.

ProblemTaskPointsFile
1Implement the reward model's ranking loss and train on preference data40src/reward_model.py
2Complete VERLTrainer and run RLHF training40src/rlhf_trainer.py
3Evaluate and analyze the model before and after RLHF20run scripts/evaluate.py

The starter code marks what you need to fill in with BEGIN ASSIGN7_* / END ASSIGN7_*. There are four spots:

  • compute_loss: the reward model's ranking loss (Problem 1)
  • _compute_gae: advantages via GAE (Problem 2.1)
  • the PPO loss inside _train_step_custom (Problem 2.2)
  • _compute_entropy: entropy from logits (Problem 2.3)

The defaults in src/config.py tell you the scale. The policy is gpt2, the reward model is distilbert-base-uncased, PPO clip is 0.2, GAE λ is 0.95, and every batch size is in single digits.

Grading criteria

The page ends with three rules:

  1. Problem 1: pass pytest and reach at least 60% validation accuracy on the reward model.
  2. Problem 2: show a reasonable reward increase in rlhf_training_curves.png. The page's example is from -0.5 to +0.5.
  3. Problem 3: show clearly different reward distributions before and after RLHF in reward_comparison.png, and upload checkpoints of the best reward model and best RLHF model so TAs can reproduce the evaluation.

Problem 3 also carries a warning. You may see gibberish generations getting high rewards. The page says this is expected: the reward model is a neural network, and the policy learns to exploit its loopholes. That's reward hacking, and seeing it in your own run teaches more than reading about it.

What "VERL-like" actually means

This is the easiest thing to misread in the starter code. requirements.txt lists verl>=0.1.0, and the classes are named VERLPolicyWrapper, VERLValueWrapper, and VERLTrainer. But src/rlhf_trainer.py imports only PyTorch and transformers. It never imports verl.

So Assignment 7 trains you on the RLHF dataflow: generate rollouts, score them with the reward model, compute advantages, update with PPO. It doesn't touch the distributed problems HybridFlow actually solves, such as coordinating multiple controllers or resharding the model between training and generation. For that layer, read VERL's source and docs.

Seen across the whole course, this makes sense. In HW5 you wrote data and pipeline parallelism yourself. In HW6 you used DeepSpeed and SGLang. HW7 shrinks the scale to GPT-2 so you can focus on the RLHF pipeline itself.

Timeline and versions

  • Deadline: the Syllabus puts "HW7 Due" on the 4/20 row (11/30 in Fall 2026). It doesn't list a release date.
  • Order relative to the lecture: the RL systems lecture is on 4/15, five days before the deadline. In practice you start the assignment before the lecture, and the algorithm background comes from the assignment's readings.
  • Repo history: commits in December 2025 fixed the GAE reward shift and added a KL penalty and the grading criteria. On 2026-03-18 a commit fixed the KL approximation and prompt/response alignment in evaluation. On 2026-05-02 a pull request was merged.
  • README vs. assignment page: the repo README's header says "Spring 2025," and its clone command uses llmsys_f25_hw7, which GitHub redirects to llmsys_hw7. The assignment page clones llmsys_hw7 directly. Go with the assignment page.
  • Required or optional: Logistics says only that 2 of the 7 assignments are optional, not which ones. The official pages don't say whether HW7 is required.

Where self-learners get stuck

  • Hardware: the page states no hardware requirements. GPT-2 and DistilBERT are far lighter than HW5 and HW6, but there are no official numbers for training time or GPU memory.
  • Grading: the only public pytest file is tests/test_reward_model.py, which covers Problem 1. Problems 2 and 3 are judged by TAs from plots and checkpoints. Off campus, nobody checks them, so grade yourself against the three criteria on the page.
  • Submission: the page asks for assignment7_[your_andrew_id].zip. Self-learners have no way to submit and no private tests.
  • Data: you need to download HH-RLHF from Hugging Face. requirements.txt also lists wandb and tensorboard; one of them is enough.

How to study it

  1. Read the RLHF primer the page links, until you can explain how chosen/rejected pairs, the reward model, and PPO fit together.
  2. Do Problem 1: get pytest passing and push validation accuracy past 60%.
  3. Before Problem 2, read how generate_rollouts in rlhf_trainer.py builds a RolloutBatch. Then fill in GAE, the PPO loss, and entropy.
  4. After Problem 3, look at a few high-reward generations and find a reward-hacking example.
  5. If you have time, read the first architecture figure in HybridFlow. Compare it with the single-machine trainer you just wrote and list which steps would become bottlenecks at scale.

One thing to do tonight: clone llmsys_hw7, open src/config.py, and match each PPO parameter to the PPO formula you know. The ones you can't place are tomorrow's reading.

Where to learn the algorithms

This course treats RLHF as a systems problem. Guides to other courses on this site cover the algorithms more fully:

References