🌏 中文版
Source term: This post is based on the Spring 2026 Homework 1 PDF, LaTeX template, and starter code hw1_starter_code.zip for CS224R. I downloaded and read all three anonymously on 2026-09-30. There is no video for this assignment.
This is part 3 of the Reading Stanford CS224R series. It follows L2 on imitation learning, which covered three ideas: why a policy needs to represent multimodal distributions, action chunking, and online interventions with DAgger. HW1 puts all three into one small game so you can see for yourself what problem each one solves.
The course schedule says HW1 went out on Friday, April 3, the day of L2, and was due April 10 at 9 pm Pacific. It is worth 10% of the course grade.
How much is public
On this site's course map scale, this assignment is A3 (enough to self-study). The problems, template, and full starter code are public, and you don't need cloud compute.
What you can't get:
- Solutions, the autograder, and the Gradescope submission pages
- Clarifications and errata posted on Ed
- TA office hours
Also note the AI-tool rule in the HW1 PDF itself. To make sure students really understand how these imitation learning methods are implemented, the assignment prohibits using generative models to help write code. That is stricter than the general policy on the course homepage, which allows discussing problems with AI tools but requires you to write solutions independently. Self-learners aren't bound by the honor code, but the rule tells you what the assignment is for. None of the three algorithms is long, and you learn from writing them yourself.
The environment: a bird driven by a PD controller
This isn't the original Flappy Bird. Here is how the PDF and the starter code README describe it:
- Action: one number between 0 and 1, the bird's target height. Inside the environment, a PD controller turns the target into thrust, so the bird has momentum and the policy has to anticipate.
- Observation: 4-D and normalized. It holds the horizontal distance to the next pipe, the height of gap 1, the height of gap 2, and the bird's current height. In easy mode, gap 2 equals gap 1.
- Easy mode: each pipe has one opening.
- Hard mode: single-opening and double-opening pipes alternate.
- Success: survive all 1000 steps. Hitting a pipe or the screen edge ends the episode.
The demonstrations come from Expert in expert.py, which is read-only. Its docstring spells out how it behaves in hard mode. While the pipe is still far away, it hovers midway between the two openings. Once it gets within a set distance (commit_dist, default 0.18), it randomly picks one opening and targets it until the next pipe appears. Its actions are also smoothed with an EMA.
Read all of expert.py before you start. What happens in the three problems traces back to how this expert generates data.
Action chunking: predict 20, execute 10
The PDF says the policy outputs ACTION_CHUNK=20 future target heights at once. During a rollout, only the first EXECUTE_STEPS=10 run before the policy is queried again. The PDF calls this receding horizon control. It notes that most robot learning policies today work this way and that it usually performs much better.
The L2 slides cover the same idea. They list papers that introduced action chunking, including Diffusion Policy, which is also an assigned reading on the schedule.
In code, this means the BC policy's output has 20 dimensions, not 1. In networks.py, BCPolicy defaults to state_dim=4, action_dim=20, hidden=256.
The code you write: seven TODOs
Every TODO raises NotImplementedError. The PDF suggests this order:
| Order | File::function | Problem | Spec from the PDF |
|---|---|---|---|
| 1 | networks.py::BCPolicy | Problem 1 | 3-layer MLP: Linear → ReLU → Linear → ReLU → Linear → Sigmoid |
| 2 | BC loss in losses.py | Problem 1 | MSE between predicted and expert actions |
| 3 | networks.py::FlowMatchingSchedule.interpolate, .sample | Problem 2 | See the flow matching section below |
| 4 | losses.py::flow_matching_loss | Problem 2 | Tip: call schedule.interpolate |
| 5 | dagger.py::DeterministicExpert.act | Problem 3 | Same logic as hard-mode Expert.act, but pick a strategy that resolves the ambiguity |
| 6 | dagger.py::rollout_episode | Problem 3 | Reset the environment, use the policy's action chunks, return state-action pairs |
| 7 | dagger.py::rollout_and_relabel | Problem 3 | Roll out with rollout_episode, then relabel with DeterministicExpert |
main.py, visualization.py, flappy_bird_env.py, and expert.py are all marked read-only. The network for flow matching, a 1-D conditional U-Net called ConditionalUnet1D, is already written. You only write the schedule and the loss.
Problem 1: Regression BC (2 points)
Start with the simplest version: an MLP maps the state straight to a 20-step action chunk, trained with MSE against the expert.
Experiments and questions:
- Run
python main.py --method bc_reg --env easyand report the mean and standard deviation of episode length over 50 evaluation episodes. - Run
python main.py --method bc_reg --env hardand report the same. No new code is needed. - In 2–3 sentences, explain how MSE regression performs in hard mode and why.
Question 3 is the turning point of the whole assignment. Think back to how the expert picks an opening in hard mode, and to why L2 stressed that a policy needs to represent multimodal distributions. This guide won't write those 2–3 sentences for you. Answer with the numbers you got.
Problem 2: Flow Matching (2 points)
Intuition: learn a velocity field that flows from noise to the answer
Regression BC gives one answer per state. Flow matching instead learns a generative model. Given a state, it starts from random noise and pushes that noise, step by step, into a plausible action chunk.
Training is simple:
- Take a real action chunk from the dataset and draw Gaussian noise of the same shape.
- Draw a random time τ between 0 and 1 and take the point that far along the straight line from the noise to the real action. The closer τ is to 1, the more that point looks like the real action.
- Show the network the state, this in-between point, and τ, and have it predict which way to move. The correct answer is the direction "real action minus noise."
At inference, start from pure noise and take n small steps in the direction the network predicts. The starter code defaults to 20 steps (NUM_DIFFUSION_ITERS = 20). The result is an action chunk. Because each run starts from different noise, the same state can produce different action sequences.
The PDF says flow matching is similar to diffusion but simpler to implement, and it usually performs as well or better.
Formulas from the PDF (interpolation, loss, Euler integration)
Let $a_t$ be an action chunk from the demonstrations and $a_{t,0} \sim \mathcal{N}(0, I)$ noise of the same shape. Sample $\tau \sim U(0,1)$ and define the interpolation
$$a_{t,\tau} = \tau a_t + (1-\tau) a_{t,0}$$
Train a network $v_\theta$ to predict the velocity that moves $a_{t,\tau}$ toward $a_t$:
$$\mathcal{L}{FM}(\theta) = \frac{1}{|\mathcal{D}|}\sum{(s_t, a_t)\in\mathcal{D}} \left| v_\theta(s_t, a_{t,\tau}, \tau) - (a_t - a_{t,0}) \right|_2^2$$
At inference, start from $a_{t,0} \sim \mathcal{N}(0,I)$ and integrate $\frac{da_{t,\tau}}{d\tau} = v_\theta(s_t, a_{t,\tau}, \tau)$ from $\tau=0$ to $\tau=1$. The simplest method is Euler:
$$a_{t,\tau+\frac{1}{n}} = a_{t,\tau} + \frac{1}{n} v_\theta(s_t, a_{t,\tau}, \tau)$$
Repeat n times to get $a_{t,1}$, the action chunk that gets executed. The PDF says sample must clamp its result to [0, 1]. The starter code docstring calls this schedule conditional optimal-transport flow matching.
What to implement and answer
FlowMatchingSchedule.interpolate: given a clean action chunk and τ, sample noise and return the interpolated point and the target velocity.FlowMatchingSchedule.sample: start from Gaussian noise, runnum_stepsEuler steps, and clamp the result to [0, 1].flow_matching_loss: implement the loss above.- Run
python main.py --method bc_flow --env hard, report the mean and standard deviation, and explain its hard-mode performance in 2–3 sentences.
If you want the math behind flow matching from the ground up, this site's MIT 6.S184 flow matching lecture guide starts from ODEs. For this assignment, the intuition above and the three equations in the fold are enough.
Problem 3: DAgger (2 points)
Problem 3 takes a different route. It keeps the model, the MSE regression policy from Problem 1, and changes the data instead.
Here is how the PDF describes it. Repeatedly roll out the current policy, collect the states it visits, and have the expert relabel each state with its action, giving $\mathcal{D}{DAgger} = {(s, \pi{expert}(s)) \mid s \sim \mathcal{D}_\pi}$. Merge this with the original data and retrain with the same regression objective. This eases the distribution shift between the expert and the learned policy. The policy reaches states the demonstrations never covered, and DAgger fills in expert labels there.
The expert here is not the original Expert but the DeterministicExpert you write. The PDF asks for the same logic as hard-mode Expert.act, but with a strategy that resolves the ambiguity from the earlier problems, so that the MSE regression policy can succeed.
Experiments and questions:
- Run
python main.py --method dagger --env hardwith the default 5 rounds. Plot a learning curve with the round on the x-axis and mean episode length with standard-deviation error bars on the y-axis. Draw the Problem 1 regression result as a horizontal line on the same plot. - Comparison (0.5 points): compare regression, flow matching, and DAgger (final round) in hard mode with a bar chart or table.
python main.py --plotplots the latest run inresults/. - Answer in 3–4 sentences (0.5 points): why does DAgger improve over rounds? What role does the deterministic expert play? How does this approach fix the problem MSE regression ran into earlier?
Problems 2 and 3 are two different fixes. One makes the model able to represent several answers. The other makes the data contain only one. Putting both results on one chart is where this assignment most rewards your time.
What you submit
- Written: a PDF report with results for Problems 1–3, submitted to "Homework 1 (Written Part)" on Gradescope.
- Code: a zip submitted to "Homework 1 (Programming Part)". It holds the
hw1/folder (with the TODOs innetworks.py,losses.py,expert.py, anddagger.pyfilled in) plus four result files:bc_reg_easy.txt,bc_reg_hard.txt,bc_flow_hard.txt, anddagger_hard.txt.
Training settings in the starter code
main.py shows the settings that actually run. They affect how you read your results, so it helps to know them up front:
| Item | Setting |
|---|---|
| Expert demos | 500 episodes per mode, cut into action-chunk training data |
| Regression BC | 100 epochs, learning rate 1e-5, batch size 2048 |
| Flow matching | 50 epochs, batch size 2048, 20 integration steps at inference |
| BC and flow evaluation | 50 episodes |
| DAgger | 5 rounds, 30 rollout episodes per round, 50 evaluation episodes per round; a final 100-episode evaluation is saved to the result file |
The device is picked automatically in the order CUDA → MPS (Apple Silicon) → CPU. The README says a GPU isn't required but speeds training up a lot. colab_instructions.md gives steps for Colab with a T4 GPU.
File mismatches you'll hit when self-studying
These are small gaps between the PDF and the starter code. They don't change the assignment, but they are confusing on a first read:
- The BC loss name: the PDF says
bc_loss, whilelosses.pyand the README call the functionmse_loss. Use the name in the code. - Problem numbers: code comments label the flow matching loss "Problem 3" and the DAgger TODOs "Problem 4". The PDF has only three problems: flow matching is Problem 2 and DAgger is Problem 3. Go by the PDF.
- References to things that aren't there: comments in
networks.pyandlosses.pytell you to "compare withDDPMSchedule" and "compare withdiffusion_loss", but neither exists in this version of the starter code. - requirements.txt: the README's folder map lists
requirements.txt, but the zip doesn't include it. Install with the pip command ininstallation.mdinstead, which liststorch gymnasium pygame matplotlib "imageio[ffmpeg]" "numpy==2.2.4". - Too many hints: the docstrings at the top of
main.pyanddagger.pystate outright what happens in Problem 1 and howDeterministicExpertshould be designed. If you want to work it out yourself, read the PDF andexpert.pyfirst and those two file headers last. - A typo: the PDF's hard-mode description spells double as "doulbe".
Something to do tonight
Download the starter code, set up the environment with installation.md, write only BCPolicy and the BC loss, and run easy mode. It runs on a CPU. Once you have your first numbers, switch to hard mode with the same command and see how far the result drops. Both later problems start from that gap.
Further reading
- L2: Imitation Learning: the theory behind this assignment
- Berkeley CS285: Imitation Learning and RL Basics: another course's take on DAgger through distribution shift
- MIT 6.S184 Flow Matching: the full flow matching derivation
- CME295 Diffusion LLMs: diffusion-style methods applied to language models
Series navigation: Previous L2: Imitation Learning | Next L3: Policy Gradients | Series overview
References
- CS224R: Deep Reinforcement Learning (Spring 2026 homepage and schedule)
- CS224R Spring 2026 Homework 1 PDF
- Homework 1 LaTeX template
- hw1_starter_code.zip
- L2 Imitation Learning slides (2026)
- Chi et al. 2024: Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Loading...