Skip to content

CMU 10-423 L20: Reasoning Models — From Chain-of-Thought to o1, DeepSeek-R1, and GRPO, Plus a Look at Mechanistic Interpretability

Sep 30, 20261 min
TL;DRLecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.

🌏 中文版

This post is based on the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is part 19 of the Reading CMU 10-423 series and follows L19 + L21: long context and state space / hybrid models. It covers Lecture 20, "Reasoning Models," on March 30, 2026, given by Aran Nayebi and Matt Gormley.

Official materials used: the schedule and the L20 slides (39 pages, no inked version). The full title on the cover slide is "Reasoning Models + Mechanistic Interpretability," which adds the interpretability second half that the schedule leaves out. The schedule lists no readings for this lecture, so this post cites only the slides and the sources they credit. The course's access grade is A3 (definitions in the global AI/CS course map), but the recordings sit behind a CMU Panopto login, so this post relies only on the slides.

The question this lecture answers: how do reasoning models differ from ordinary LLMs in training and inference? The slides answer in three steps: get the model to write its reasoning out, reward correct reasoning with reinforcement learning, and let the model spend more compute thinking at inference time too.

One thing first: the exam is tonight

The second slide is a reminder: an 80-minute exam at 7 pm that evening, covering Lectures 1–15 (the same as Quizzes 1–4), with one double-sided sheet of notes allowed; unlike the all-multiple-choice quizzes, the exam includes open-ended questions. So L20 itself is not on the exam. It is tested only by Quiz 5 (April 6, covering L16–L20), whose questions are not public.

Step one: get the model to write its reasoning out

Chain-of-thought prompting

The slides pick up from in-context learning in L10:

A really long "thought"

The slides then spend eight pages on the cipher problem from OpenAI's o1 announcement. The prompt gives an example: oyfjdnisdr rtqwainr acxz mynzbhhx decodes to Think step by step, and asks the model to decode another ciphertext by the same rule.

The excerpted thinking reads like someone trying things on scratch paper: count the letters, notice that each ciphertext word is exactly twice as long as the plaintext word, guess that two letters map to one, try summing letter positions, and keep going until it decodes THERE ARE THREE R'S IN STRAWBERRY. The slides mark "1276 lines later" in the middle, a reminder of how long this thinking runs.

The point is not the cipher. It is a product decision behind o1: OpenAI did not release the full "Thinking" output, only a summary of it.

STaR: train on the reasoning that got it right

Before o1, the slides introduce STaR (Self-Taught Reasoner). There are only two kinds of data: a small set of human-annotated rationales, and many problems without rationales. The loop repeats:

  1. Use the few rationale examples for ICL to generate rationales for the problems without them
  2. If a generated answer is wrong, try to regenerate a rationale that leads to the correct answer
  3. Fine-tune on all rationales that led to correct answers

This step turns "reasoning" from a prompting trick into training data.

Step two: train thinking tokens with reinforcement learning

AIME and o1

The slides first introduce the benchmark that keeps coming back: the AIME 2024 dataset, problems from the American Invitational Mathematics Examination.

Then the summary of o1:

  • o1 was trained with reinforcement learning to generate chain-of-thought style rationales for its answers
  • These rationales (called Thinking tokens) are hidden from the user, who sees a summary instead
  • At train time, compute can be increased by doing more reinforcement learning; at test time, by letting the model think longer
  • Result 1: more train-time compute gives higher accuracy on reasoning problems; result 2: more test-time compute does too

The slides ask and answer their own question here: why is this description so vague and non-technical? Because OpenAI only released a blog post, and this is about the sum total of what it said.

The slides then use OpenAI's figure to show o1 beating GPT-4o across math, reasoning, commonsense, coding, and other problems, and conclude: the closed-source o1 was clearly superior to any open-source model, "so we waited for the open source models to catch up…"

Enter DeepSeek-R1

DeepSeek-R1 is the slides' answer: open source, open weights, 671B parameters, a carefully tuned version of the base model DeepSeek-V3, with performance comparable to o1.

How PPO and GRPO differ

To explain how R1 was trained, the slides go back to the algorithm. GRPO predates R1 and was introduced in DeepSeekMath. The slides' one-line version: GRPO is an RL algorithm akin to PPO, but it removes the value model and so greatly reduces memory requirements.

The slides paste two passages and the diagram from the DeepSeekMath paper. The differences fit in one table:

PPOGRPO
Models to trainPolicy model + value modelPolicy model only
Where the advantage comes fromRewards and the value model's estimates, via GAESample a group of answers (o₁…o_G) to the same question and use their relative rewards as the baseline
Where the KL penalty goesAdded to the per-token rewardAdded directly to the loss, keeping the advantage computation simple

For PPO, look back at RLHF in L11: the PPO there is the left column of this table.

Expand: the structure of the GRPO objective

Equation (3) from DeepSeekMath, as pasted on the slides, breaks into three layers:

  1. For each question q, sample G answers from the old policy
  2. For every token of every answer, compute the probability ratio between the new and old policies, multiply by the advantage Â, and apply the same clip as PPO (range 1−ε to 1+ε)
  3. Average over all answers and tokens, then subtract β times the KL between the new policy and the reference policy

ε and β are hyperparameters. The only differences from PPO are where  comes from and where the KL goes.

DeepSeek-R1-Zero: reinforcement learning only

The slides spend five pages on R1-Zero:

Training method

  • Trained entirely with reinforcement learning, without any supervised fine-tuning (SFT)
  • Starts from the pretrained DeepSeek-V3-Base, with RL that uses no human preferences
  • The slides call it one of the first large-scale demonstrations of RL-only training for LLMs
  • The aim was to see whether reasoning abilities can emerge from RL alone, without labeled data

Reward model: no neural reward model, just two rule-based rewards:

  • Accuracy reward: is the answer correct?
  • Format reward: did the model follow the prompt template?

The template asks the model to put its reasoning inside <think> tags and its answer inside <answer> tags.

Results

  • On AIME, the longer the model is trained with RL, the better it performs, eventually surpassing o1
  • The model gradually learns to use longer and longer sequences of Thinking tokens. This comes purely from the RL objective; nothing directly pushes reasoning length up

Problems

  • Poor readability: humans can't really understand what it is saying
  • Language mixing: English and Chinese muddled into a pidgin

DeepSeek-R1: bring SFT back

R1 builds on R1-Zero with a hybrid training strategy. The figure the slides cite lists four steps:

  1. Cold start: fine-tune the base model on a few thousand curated, human-friendly long CoTs
  2. Reasoning-focused RL: scale up RL on math, coding, and logic tasks, and add language-consistency rewards to keep the model in a single language
  3. Rejection sampling + SFT: sample correct, well-structured chains of thought from the RL model, add general-capability data (writing, Q&A, self-cognition), and train a new base checkpoint
  4. RL across scenarios: a second RL stage covering both reasoning tasks and general tasks, for "helpfulness" and "harmlessness"

The slide text calls this a "two-stage pipeline," while the cited figure lists four steps. My reading is that the "SFT then RL" pair is done twice, but that is a guess; the in-class explanation is not available. The slides' conclusion is clear: doing SFT before RL fixed R1-Zero's repetition and language mixing, and improved readability, coherence, and task accuracy.

Second half: mechanistic interpretability

The last six slides switch topics. They list four reasons interpretability matters: safety (corrective action), safety (preventative action), preventing an AI apocalypse, and learning from AI.

Why it is hard: the main problem is superposition. A human-interpretable "feature" rarely activates in a single place in the network; its activations are almost always spread across many locations: across heads, across MLP neurons, across layers.

Replacement models: for specific blocks in a network, train a "replacement block" to mimic the original block's input-to-output behavior, with the key requirement that the replacement be more interpretable. The techniques listed:

  • Sparse autoencoders: replace the MLP layer in a Transformer block with an autoencoder version that has more neurons (features) in the hidden layer, plus a regularizer that encourages sparse activations, such as L1
  • Circuits
  • Cross-layer transcoders: let a replacement block directly access all earlier replacement blocks

The final slide shows Anthropic's On the Biology of a Large Language Model, which uses a circuit tracing method to study the internal mechanisms of Claude 3.5 Haiku in settings such as multi-step reasoning, planning rhymes, multilingual circuits, medical diagnoses, refusals, jailbreaks, and CoT faithfulness. The slides show only the study's table-of-contents figure and do not walk through individual cases.

One table to wrap up

StageExamplesWhere the reasoning comes fromWhere compute is added
PromptingCoT, zero-shot CoTDemonstrations, or a single "think step by step"A few more tokens at inference
Self-trainingSTaRGenerated by the model, keeping only correct onesFine-tuning
Reinforcement learningo1, R1-Zero, R1Rewards for correct answers (R1 adds format and language consistency)More RL at training time, longer thinking at inference time

Try this tonight: compute GRPO advantages by hand once. Suppose you sample 4 answers to one question and the rule-based rewards are [1, 0, 0, 1]. Following the outcome supervision setup in the DeepSeekMath paper, subtract the group mean (0.5) and divide by the group standard deviation (0.5 using the population standard deviation), giving advantages [1, −1, −1, 1]. Now change the rewards to [1, 1, 1, 1]: after subtracting the mean everything is 0, and the formula would divide by 0. The small example shows that when a group is all right or all wrong, there is nothing to compare against inside the group, so that question provides no learning signal.

What this post can and cannot confirm

Confirmed: the schedule's date and quiz coverage; the text, tables, figure captions, and credits on the slides; how GRPO normalizes advantages (checked against the DeepSeekMath paper); and the titles of the papers the slides cite (checked against arXiv). Not confirmed: what was said in class (Panopto requires login), numbers shown only in figures (for example, o1's compute curves, R1-Zero's AIME accuracy curve, and the R1 versus o1 benchmark bars), and the official explanation of the gap between "two-stage" and the four-step figure.

Further reading: this site's CME295 LLM reasoning post and CS336 RLVR post cover reasoning models and verifiable rewards from other angles; for interpretability, continue with the CS224N interpretability post and the Harvard CS2881R interpretability post.

Series navigation: previous L19 + L21: long context and state space / hybrid models | next L22 + L26: practical risks and the science of alignment | series overview

References