Skip to content

CS224R L1: Framing Decision-Making as an RL Problem

Sep 30, 20261 min
TL;DRThe first lecture of CS224R Spring 2026 does three things: covers logistics, explains why deep RL is worth learning, and turns 'behavior' into something you can learn. The core is a set of definitions (state, action, trajectory, reward, policy) and one objective: maximize expected total reward. It ends on an example: fit ℓ2 regression to drivers where some change lanes and some go straight, and the policy learns their average, a half lane change nobody demonstrated. That problem is where L2 starts.

🌏 中文版

Source year: based on the Spring 2026 01_cs224r_intro_2026 slides (2026-04-01). The companion video is the Spring 2025 L1 recording (supplement). The title matches, but the slides were revised for 2026, so details may differ. This is post 1 of the Reading Stanford CS224R series.

On the CS224R schedule, lecture 1 is called "Course Intro + Start of MDPs & Imitation". The slides list three learning goals: how to represent behavior, how to formulate a reinforcement learning problem, and the basics of imitation learning.

The logistics (grading, late days, the AI tools policy) are covered in the series overview. This post covers only the content.

First, what an MDP is

The official prerequisites assume some familiarity with RL, and the slides say MDPs will get a quick pass. If you have never seen one, this intuition is enough to start:

An agent sees the state of the world, picks an action, the world moves to a new state as a result, and the agent gets a reward. Then it repeats.

An MDP writes that loop down as math. Its key assumption is that the next step depends only on the present. Given the current state and action, you know the distribution over the next state, with no need for anything earlier.

For fuller background, the course recommends two entry points: chapters 3–4 of Sutton & Barto, or the MDP and RL modules from CS221. On this site, CS221 L7: MDPs I and L8: MDPs II cover exactly this.

What deep RL means

The slides define it in two halves. The problems are sequential decision-making problems: a system makes many decisions based on a stream of information, observing, acting, observing again, acting again. The solutions include imitation learning, offline and online RL, RL for LLMs, model-free and model-based RL, multi-task and meta RL, and RL for robots. "Deep" means the emphasis is on methods that scale to deep neural networks.

One comparison table on the slides shows how it differs from supervised learning:

Supervised learningReinforcement learning
What you learnGiven labeled data {(xᵢ, yᵢ)}, learn f(x) ≈ yLearn behavior π(a | s)
FeedbackTold directly what to outputFrom experience, indirect
Data distributionInputs x are i.i.d.Not i.i.d.: actions affect future observations

The last row is the one the course keeps returning to. In RL, the model's outputs change the data it sees next. Compounding errors in L2 and offline RL in L7 both grow out of this.

The slides' examples of behavior are motor control, chatbots, game playing, driving and web agents.

Why study it

The slides give four reasons:

  1. Going beyond (x, y) supervision. Model predictions have consequences. When direct supervision is not available, RL can learn from any objective, including ones that are not differentiable and not just accuracy. Systems that interact with people (chatbots, recommenders) and systems whose deployment changes future observations (the slides call these feedback loops) fall here.
  2. It is widely deployed. The examples are legged robots, robot manipulation, "Move 37" in AlphaGo's match against Lee Sedol, traffic control, making image generation models follow prompts, and Google's RL for TPU chip design. The slides add that nearly all modern language models use some form of RL in post-training, especially for advanced reasoning.
  3. Learning from experience looks fundamental to intelligence. The example is robot research by Levine, Finn and colleagues from 2015–2016: robots getting better with practice, first "with its eyes closed", then with vision.
  4. There are plenty of open research problems.

The slide for the fourth point is the most useful, because it maps each research question to a later lecture:

Research questionLecture
How does a robot learn what is good or bad for the task?Reward learning (L8)
How do you use large, diverse datasets?Offline RL (L7)
How do you transfer from other tasks and goals?Multi-task RL, meta-RL (L12–L13)
Can RL learn long-horizon tasks like cooking a meal?Hierarchy, reasoning (L15, L10)
Can robots practice fully autonomously?RL for robots (L16–L17)

Between these sits a photo titled "Behind the scenes of RL": a robot practicing, with an arrow pointing to a person in the background, Yevgen. The caption reads "Yevgen is doing more work than the robot!" and adds that collecting lots of data this way is not practical. That photo is why later lectures cover offline RL and autonomous practice.

Turning experience into data

This is the heart of the lecture. The slides start with four definitions:

  • state sₜ: the state of the world at time t
  • action aₜ: the decision taken at time t
  • trajectory τ: the sequence of states and actions (s₁, a₁, …, s_T, a_T), which can have length 1
  • reward r(s, a): how good this s, a is

Add unknown dynamics p(sₜ₊₁ | sₜ, aₜ). The next state is a function of only the current state and action (plus randomness), independent of sₜ₋₁. That is the Markov property.

What if you cannot see the full state? The slides give two options: treat sensor readings as an approximation (camera images with good visibility are often close enough), or model partial observability explicitly with an observation oₜ. The cost is that observations are not Markov. Once you marginalize out the states, the next observation depends on all past observations: p(oₜ₊₁ | o₁:ₜ, a₁:ₜ) ≠ p(oₜ₊₁ | oₜ, aₜ).

Two examples on the slides make the definitions concrete:

Robot hanging a towelChatbot
State / observationstate: RGB images, joint positions and velocitiesobservation: the user's latest message
Actioncommanded next joint positionthe chatbot's next message
Trajectory10 seconds at 20 Hz, T = 200a conversation of variable length
Reward1 if the towel is on the hook, else 01 for an upvote, -10 for a downvote, 0 for no feedback

The class then does a think-pair-share: define the state, action, trajectory and reward for autonomous driving, a web agent, or a poker player. It is worth actually doing, and the exercise at the end of this post builds on it.

Representing behavior with a neural network

A policy πθ(a | s) is a neural network that takes a state and outputs a distribution over actions. At run time you observe sₜ, sample an action aₜ from πθ(· | sₜ), and the world produces sₜ₊₁ from its unknown dynamics. Repeat, and the resulting trajectory is also called a roll-out or an episode.

If you only have observations o, the slides suggest giving the policy memory: πθ(aₜ | oₜ₋ₘ, …, oₜ).

The RL objective: expected total reward

The obvious objective is to maximize the sum of rewards Σₜ r(sₜ, aₜ). But that quantity is not deterministic. The slides ask where the variability comes from. Two places: the world is stochastic, and the same policy may not make the same decision every time.

So a trajectory is itself a distribution:

pθ(τ) = p(s₁) · ∏ₜ πθ(aₜ | sₜ) · p(sₜ₊₁ | sₜ, aₜ)

and the RL objective is to maximize the expected total reward:

max_θ  E_{τ ~ pθ(τ)} [ Σₜ r(sₜ, aₜ) ]

Why stochastic policies? The slides give two reasons. To learn from your own experience you have to try different things (exploration). And existing data already shows varied behavior. The second point sets something up: we can borrow tools from generative modeling and treat the policy as a generative model of actions given states. All of L2 is about that.

How good is a policy? Two functions:

  • value function V^π(s): expected future reward starting at s and following π
  • Q-function Q^π(s, a): expected future reward starting at s, taking a, then following π

Five algorithm families, five sets of trade-offs

For the same objective, the slides list five kinds of solution. Together they are the table of contents for the first half of the course:

FamilyIdeaIn this series
Imitation learningMimic a policy that gets high rewardL2
Policy gradientsDifferentiate the objective directlyL3
Actor-criticEstimate the current policy's value and use it to improve the policyL4, L5
Value-basedEstimate the optimal policy's valueL6
Model-basedLearn a dynamics model and use it for planning or policy improvementL11

Why so many? The slides' answer: each makes different trade-offs and works best under different assumptions. The questions to ask:

  • How easy and cheap is it to collect data with the policy? (A simulator, or by hand?)
  • Which supervision is cheaper: demonstrations or detailed rewards?
  • How much do stability and ease of use matter?
  • How high-dimensional is the action space? Continuous or discrete?
  • Is the dynamics model easy to learn?

These five questions are a good lens for every later lecture.

Imitation learning version 0: why it learns the mean

The last dozen slides start on imitation learning, which they describe as both a subroutine in some RL algorithms and a strong approach on its own.

The setup: given expert demonstrations (from some unknown π_expert), learn a πθ that performs as well as the expert. The example is a dataset of human drivers, sensor readings plus steering commands.

Version 0 is the most direct: a deterministic policy, trained by supervised regression on the expert's actions to minimize ‖a − â‖², where â = πθ(s). Then deploy it.

The slides then ask what a policy trained with ℓ2 regression will do. The picture is a highway, and the demonstrations contain two steering commands: some drivers merge left (around -2) and others stay straight (around 0). Together that is two peaks, but ℓ2 regression learns the mean of the data, around -0.5. That value sits between the peaks: a half-merge nobody demonstrated. The slides say this happens "All the time!", especially when data is collected by multiple people.

So the question becomes how to represent more than the mean. The slides give two starting points:

  • Discrete actions: the network outputs a probability for each action, a categorical distribution, which is maximally expressive
  • Continuous actions: the network outputs μ and σ, a Gaussian, which is not very expressive

A Gaussian has one peak and cannot represent "left" and "straight" at once. How to represent multimodal continuous distributions with a neural network is the next lecture's topic.

What you can do tonight

Pick a system you know well (a support bot, a recommender, an agent you wrote) and write four lines, following L1's think-pair-share:

state or observation:
action:
trajectory length:
reward:

Then ask two questions. Is your observation Markov? If not, how much history does the policy need? And if you had a batch of human demonstrations where the same situation has two reasonable responses, what would ℓ2 regression learn?

Further reading

Series navigation: Previous: Series overview | Next: L2: Imitation learning and policies that can represent multimodal distributions

References