Reading Stanford CS234 Reinforcement Learning through the Winter 2026 slides for 14 lectures and three assignments with starter code, alongside the public Spring 2024 videos: MDP planning, model-free evaluation and control, policy gradients and PPO, imitation learning and RLHF/DPO, exploration theory with bandits, MCTS, and value alignment.
CS234 is Emma Brunskill's introductory reinforcement learning course at Stanford. It runs from planning in known MDPs through policy gradients, RLHF/DPO, bandit exploration, and MCTS. For Winter 2026, all 14 slide decks, the three assignment handouts with starter code, and the project spec can be downloaded without logging in, so this series rates it A3 (enough to self-study). The gaps: the site links no 2026 recordings, L15 and L16 have no slides, and the midterm and tutorials are not public. The public recordings are from Spring 2024, and this series uses them only as a supplement. Two 2024 lectures on offline RL have no counterpart in the 2026 slides.
Lecture 1 of CS234 Winter 2026 first answers what RL is: learning from experience to make good decisions under uncertainty. It usually involves four things at once: optimization, delayed consequences, exploration, and generalization. A seven-cell Mars rover world then builds from a Markov process to a Markov reward process, defining return, the value function, and the discount factor, and ends with the Bellman equation for an MRP. You can solve it with a matrix inverse or iterate with dynamic programming. Add actions and you get an MDP, where the next lecture starts.
Lecture 2 of CS234 Winter 2026 assumes the world model is known and asks how to compute the best policy. An MDP plus a policy is an MRP, so a policy can be evaluated by iterating a Bellman backup. Policy iteration alternates evaluation and improvement, and the slides prove each round is no worse than the last, so it stops within |A|^|S| rounds. Value iteration takes another route: apply the Bellman optimality operator over and over. For γ < 1 that operator is a contraction, so value iteration always converges. The lecture ends with finite horizons, where the best policy usually depends on how many steps remain.
CS234's Winter 2026 Assignment 1 is worth 68 points across four questions: an inventory MDP where the horizon and discount change the optimal policy (8), a traffic example where a proxy reward makes the AI car refuse to merge (5), bounding a greedy policy's performance with the Bellman residual (30), and hand-written value iteration and policy iteration on RiverSwim (25). The three written questions all drill one idea: the reward, γ, and value function you write down may not be the goal you think they are.
If you don't know the transition probabilities or rewards, how do you estimate what a policy is worth? CS234 Lecture 3 gives three answers. Monte Carlo averages full-trajectory returns: unbiased, high variance, and it has to wait for the episode to end. TD(0) targets one real reward plus the next state's estimate: biased, lower variance, and it updates every step. Certainty equivalence estimates a model and then runs dynamic programming: the most data-efficient and the most expensive to compute. The AB example at the start of Lecture 4 makes the difference plain: on the same data, MC says V(A)=0 and TD says V(A)=0.75.
Once you can evaluate a policy, the next step is to improve it while you collect data. CS234 Lecture 4 goes like this: ε-greedy keeps policy improvement monotonic; GLIE says how much to explore and when to stop; Q-learning converges to Q* under GLIE plus Robbins–Monro step sizes; and finally the table becomes a parameterized Q̂(s,a;w) trained by SGD on MC, SARSA, or Q-learning targets. The price is the deadly triad: function approximation, bootstrapping, and off-policy learning together can oscillate or diverge.
Q-learning converges with a table but can diverge once you add function approximation. CS234 blames the deadly triad: bootstrapping, function approximation, and off-policy learning all at once. DQN holds things together with two tricks. Experience replay breaks the correlation between consecutive samples, and fixed Q-targets keep the target still for C steps. In the Atari ablation table the slides show, Breakout goes from 3 with a linear model and 3 with a plain deep network to 317 with both tricks; replay alone reaches 241.
Policy gradients skip learning a value function and deriving a policy from it. They run gradient ascent directly on the policy parameters θ. The key step rewrites ∇P(τ;θ) as P(τ;θ)∇log P(τ;θ); after taking the log, the dynamics model drops out and only the policy's own score function is left. The raw estimator is unbiased but very noisy, and CS234 reduces the noise in three ways: pair each action only with the return that follows it (REINFORCE), subtract a state-dependent baseline (proven not to add bias), and replace Monte Carlo returns with values estimated by a critic (actor-critic).
Vanilla policy gradients have two flaws. Each batch is thrown away after one step, and distance in parameter space is not distance in policy space, so a large step can collapse performance. Following Joshua Achiam's slides, CS234 starts from the performance difference lemma, rewrites the new policy's performance as a surrogate objective over the old policy's data, and bounds the approximation error with KL divergence. Maximizing 'surrogate minus a KL penalty' guarantees no regression, but the theoretical constant is too large, so PPO approximates it with an adaptive KL penalty or clipping. Advantages come from GAE, which trades off bias and variance.
CS234 Winter 2026 Assignment 2 is worth 102 points across four questions: DQN written questions (8); REINFORCE, a neural-network baseline, and clipped PPO on three PyBullet environments, CartPole, Pendulum, and HalfCheetah (54 coding + 21 write-up); proofs about policy-induced state distributions and the performance difference lemma (14); and a Belmont Report review of an RL experiment that learns on real students (5). The coding question turns the equations from L5–L7 into code that produces 21 learning curves.
When you have expert demonstrations but no reward, the second half of CS234 L7 offers three routes. Behavioral cloning copies actions with supervised learning. DAgger fixes its compounding errors by querying the expert along the learner's own path. Inverse RL instead infers what reward the expert is optimizing. Inferring rewards runs into the fact that infinitely many rewards explain the same demonstrations; feature matching and the maximum-entropy principle are two ways to pin down an answer. This material sets up the next post on RLHF: swap demonstrations for preferences and the problem keeps almost the same shape.
CS234 L8 keeps the inverse RL problem from the previous post but changes the input. Instead of expert demonstrations, a human says "A is better than B." The Bradley-Terry model turns these pairwise comparisons into a reward you can fit with cross-entropy. RLHF runs PPO on that reward model with a KL penalty. DPO shows that the KL-constrained optimal policy has a closed form and rewrites the reward as a log-ratio of policies. Plugged back into Bradley-Terry, the partition function cancels, so you can train the policy on preference data directly, with no reward model. The 2026 slides contain no offline RL, and DPO is now taught in lecture rather than by 2024's guest speakers.
CS234 Winter 2026 Assignment 3 has five questions worth 94 points. The first three share MuJoCo Hopper: run PPO on a hand-written reward (13), learn a reward model from 10,000 preference pairs and run PPO on it (19 + 8), then learn a policy straight from preferences with SFT + DPO without ever touching the environment (6 + 19). Q4 switches to pure theory: use Hoeffding and a union bound to count how many pulls you need to find an ε-optimal arm (25). Save it until after the next post on bandits. Q5 is stated vs. revealed preferences in a news app (4).
CS234 L9 and the first half of L10 turn exploration from a rule of thumb like ε-greedy into something you can prove. First, regret: how much you lose compared with always pulling the best arm. Greedy locks onto a suboptimal arm, and ε-greedy with fixed ε spends an ε fraction of its time choosing at random, so both have regret that grows linearly with time. The Lai-Robbins lower bound says the best possible is logarithmic growth, and UCB gets there by being optimistic about uncertain arms: Theorem 7.1 of Bandit Algorithms shows each suboptimal arm is pulled only about 16 log n / Δ² times.
CS234 L11 switches the logic of exploration from optimism to sampling. Thompson sampling keeps a posterior for each arm, draws one value from each posterior at every step, and pulls the arm with the largest draw. With Bernoulli rewards and a Beta prior, the update just adds one to the success or failure count. It implements probability matching: each arm is chosen with the posterior probability that it is the best arm. Under Bayesian regret it matches UCB's order, and with batched, delayed feedback it suits the problem better than deterministic UCB. The cost: a badly wrong prior can make it perform poorly.
The previous two posts covered UCB and Thompson sampling, which only handle one-step decisions. CS234 Lecture 12 carries the same two ideas into MDPs, where states matter. First it swaps the yardstick: PAC bounds the number of steps where you act badly, not total regret. Then it covers the optimistic approach (MBIE-EB: counts plus an exploration bonus) and the sampling approach (PSRL: draw one MDP per episode and solve it). When states are too many to count, the bonus moves into the Q-learning target, which is what beat ε-greedy DQN on Montezuma's Revenge. The last section asks whether exploration itself can be learned; one answer is the Decision-Pretrained Transformer.
Until now, CS234 has computed one policy for the whole state space. Lectures 13 and 14 ask a different question: if I only care about the move in front of me, can extra local computation make that one decision better? The path runs from simple Monte Carlo search through the expectimax tree to MCTS, and treating each tree node as a bandit gives UCT. AlphaZero ties MCTS to a single network that predicts both policy and value, and self-play pushes both forward. The slides borrow figures from Silver et al. 2017 to answer three questions: how much architecture matters, how much MCTS adds, and whether human data is needed.
All of CS234 assumes the reward is given. The Winter 2026 ethics and society guest lecture (Wanheng Hu, based on material originally developed by Dan Webber) asks, over two sessions, what you really want. The first session splits "alignment" into three targets: the user's intentions, revealed preferences, and objective best interests, with RLHF-driven sycophancy and a personal AI agent as case studies. The second adds a fourth target, what is morally right for people besides the user, and compares three routes: top-down (write principles down), bottom-up (learn from examples), and participatory AI. There's no silver bullet, but alignment can be better or worse.
The last guest deck in CS234 Winter 2026, by Shane Gu of Google DeepMind: 36 slides, no public recording. It has three threads. First, Solomonoff induction says the best predictor is the shortest program that generates the data, and prediction comes in three levels. Second, a forward model F and two inverse models, Π and Q, share one notation, which shows how shooting and direct collocation each plan with a different kind of model and why TDMs and Generalized Decision Transformers are world models at a different time scale. Third, the deck asks whether video models can become the foundation model for the physical world.