Skip to content
All tags

#reasoning

25 posts

CMU 10-423 L20: Reasoning Models — From Chain-of-Thought to o1, DeepSeek-R1, and GRPO, Plus a Look at Mechanistic Interpretability

Lecture 20 of CMU 10-423 (Spring 2026) tells the story of reasoning models as one line: chain-of-thought prompting gets models to write intermediate steps, STaR fine-tunes on the reasoning that led to correct answers, and OpenAI o1 trains thinking tokens with reinforcement learning so compute can be added at both training and inference time. On the open side, DeepSeek-R1-Zero uses only rule-based rewards and GRPO and its reasoning grows longer on its own; DeepSeek-R1 adds SFT back to fix readability and language mixing. The lecture ends with mechanistic interpretability: why superposition makes models hard to read, and how replacement models such as sparse autoencoders, circuits, and cross-layer transcoders address it.

CS224R L10: RL for LLM Reasoning and Test-Time Compute

Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.

NCCU Yen-Lung Tsai Generative AI L14: Text and Image Models Invade Each Other's Territory, Plus the Final Project

The last lecture looks at two lines of technology crossing into each other. LLMs such as ChatGPT have started drawing, and the slides use early fusion plus VQ-VAE/VQGAN to explain how an image can be cut into tokens. Going the other way, Inception Labs' Mercury generates text with diffusion, noising a sentence into a row of [MASK] tokens and then restoring it. Next come a few papers anyone can use: evaluating RAG automatically, reasoning models being easier to hijack, and DeepMind's four kinds of AI risk. The lecture ends with vibe coding and a list of application tools, and the final project runs as an online conference in Gather Town.

NTHU NLP Course Summary and LLM Reasoning Notes: A Three-Column Map of the Semester, Then Denny Zhou's Question of Whether Pretrained Models Can Reason

In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.

Reading NTU ADL 2025 Fall: Reasoning — A Video-Only Lecture, Five Steps from CoT to RL

The Reasoning lecture of NTU ADL Fall 2025 has no public slides. It exists only as five videos in the course playlist: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason (imitating others), and 12.5 RL for Reasoning (evolving reasoning through exploration). This post lays out that route from the video titles alone, then pairs it with the CoT, ReAct, and 'reasoning enlarges the action space' pages of the previous Language Agents deck. Technical detail is left to the site's CS224N and CME295 reasoning posts.

Hung-yi Lee ML 2026 HW8: Spending More Inference Compute — What Voting, Self-Certainty, and DeepConf Each Buy in Accuracy

HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.

Hung-yi Lee ML 2026 Self-Correction: Can a Model Fix Its Own Mistakes? What Changing Decoding, Workflow, or Weights Buys You

This lecture asks whether a model can catch and fix its own errors with no human in the loop. Hung-yi Lee splits the approaches into three routes. Change inference: the whole contrastive decoding family builds a version of the model likely to be wrong and subtracts it, and the methods differ only in how that wrong version is made. Change the workflow: appending "check again" sometimes helps but is unstable, external feedback beats self-reflection, and under a fixed compute budget, sampling more answers and voting often wins. Change the weights: teaching self-correction directly runs into "after training, the model makes different mistakes," which is why the field moved to RL. Whether RL teaches new abilities or just makes existing paths more likely is still being debated.

CME295 Lecture 6: How Reasoning Models Learn to Think Longer, and What GRPO Drops from PPO

CME295 Lecture 6 breaks reasoning models into three pieces: emit a reasoning chain before the answer, run RL on verifiable rewards like "is the answer correct," and use GRPO, which takes the group's average reward as the baseline instead of training a value model. RL alone took DeepSeek-R1-Zero from 15.6% to 71.0% pass@1 on AIME 2024, and distilling R1's traces into Qwen-32B beat running RL on the 32B model directly.

Multi-hop Retrieval: When Answers Are Scattered Across Documents

Standard RAG retrieves one set of documents per query, but real questions often need reasoning across 2-4 documents. IRCoT pioneered interleaved retrieval-reasoning, PAR²-RAG beats IRCoT by 23.5% accuracy on four benchmarks, and CompactRAG compresses LLM calls down to just two.

Agentic / Reasoning RAG: From Search-R1's RL Multi-Turn Search to Deep Research and MCP's Reasoning × Retrieval Paradigm

In 2025 RAG stopped being 'retrieve once, generate once.' Search-R1 trains models to search autonomously in multiple turns with RL, REX-RAG/AlignRAG add policy and alignment branches, OpenAI Deep Research productizes the loop, and MCP generalizes retrieval into unified tool invocation. This post unpacks the design philosophy, trade-offs against ten generations, and when to adopt the new paradigm.

2025 AI Conference Review: Machine Learning

ML conferences broke every submission record in 2025 and pushed peer review to its limit. NeurIPS received 21,575 papers and used more than 20,000 reviewers; ICML passed 12,000 for the first time, and ICLR reached 11,565. Reasoning and agents were the strongest trends. One NeurIPS runner-up, the conference's only perfect-score paper, challenged whether RLVR creates new reasoning ability. Awards for Alibaba Qwen's Gated Attention and a mechanistic theory of neural scaling laws showed a community moving from scaling at all costs toward understanding why scaling works.

What Top AI Conferences Accepted in 2025: The Agent Breakout and Reasoning Revolution

The two strongest signals at AI conferences in 2025 were reasoning papers jumping from 47 to 216, a 4.6-fold rise, and agent-related terms exceeding 150 papers with 4.3–11-fold growth. Diffusion moved from breakout topic to infrastructure; RAG became a mainstream enterprise architecture with unusual coverage across all five conferences; state-space models and world models began tracing the early 2020–2021 path of Vision Transformers. Pure prompt-engineering papers encountered reviewer fatigue.

Gemini——Google's Native Multimodal Flagship: 1M Context and Scientific Reasoning Champion

Gemini is Google DeepMind's native multimodal LLM family, famed for a 1M-token context window and native video/speech input plus scientific reasoning. 3.1 Pro tops GPQA Diamond 94.1% and ARC-AGI-2 77.1% to claim science-reasoning dual crowns, at $2/$12—1/6 of Claude. 3.7 Flash delivers near-Pro agent capability for $0.75/$3.75.

GPT——Closed API for Revenue, Open GPT-OSS for Ecosystem: the Unified Routing Platform Behind the World's Largest AI Service

GPT is OpenAI's LLM family, from 117M parameters in 2018 to the three-tier GPT-5.6 Sol/Terra/Luna lineup in 2026, serving 1B+ users and 2M enterprise customers. GPT-5.6 Sol leads LiveBench 81.1%, Terminal-Bench 2.1 88.8%, and Artificial Analysis Coding Agent Index 80 across multiple agentic benchmarks, while OpenAI's first open-weight model GPT-OSS ships under Apache 2.0.

Kimi——From a 200K Long-Context Tool to a 2.8T Open-Source Frontier, and K3's Architectural Leap

Kimi is Moonshot AI's LLM family, born from ultra-long context. Kimi K3 (2026/07) is the world's first open 3T-class model—2.8T params, 104B active, 1M context, scoring 60 on the Artificial Analysis Intelligence Index tied with GLM-5.3 for open-source #1. Its Kimi Delta Attention brings a 2.5× scaling efficiency gain.

CS224N Lecture 19: How Small Models Can Move Beyond Brute-Force Scaling

The final lecture frames Open Questions in NLP 2026 as smart scaling: prolonged RL, Prismatic synthetic data, RL as pretraining, and open collaboration seek reasoning gains beyond adding parameters.

CS224N Lecture 12: Decoding, DeepSeek-R1, and Reasoning Training

Lecture 12 shows that output policy is not a detail: greedy, beam, and sampling produce different text. It then moves from R1-Zero/R1 into PPO, GRPO, and DAPO, asking when longer reasoning actually helps.

CS224N Lecture 13: Speculative Decoding and Test-Time Scaling

Lecture 13 moves from inference efficiency to inference capability: speculative decoding drafts with a small model and verifies with a large one; on-policy distillation addresses drift; long context and test-time scaling spend inference resources.

CS336 Lecture 16: RLVR Scales Reasoning with Verifiable Rewards, but GRPO Is Not Free PPO

Lecture 16 moves from PPO to GRPO and RLVR. Math, code, and environment outcomes provide scalable rewards and avoid some preference-model overoptimization, but group-normalized advantages introduce difficulty and length bias while rollout infrastructure becomes the dominant cost.

Berkeley CS288 Part 5: Inference-time Compute, Reasoning, and Embodied Agents

Units 15–18 place NLP models inside perception, reasoning, tool, and environment loops; the question shifts from next-token prediction to allocating inference compute and validating multi-step action.

Stanford CS329A: A Course on Self-Improvement That Says Out Loud What It Can't Improve

CS329A is built around the generation–verification gap: models can produce the right answer but can't tell which one it is. The conclusion the course draws about itself matters more — today's methods make models more consistent, not smarter. Nine lectures are public, out of twenty.

aideep-dive

Resource Rationality for Agents: Optimal Decisions Across Tokens, Tool Calls, and Latency

Agent decision-making under resource constraints is bounded rationality reborn: Rational Metareasoning uses VOC rewards to save 20-37% of tokens, BATS proves that adding budget without budget awareness is futile, FrugalGPT cascades cut costs by up to 98%, and Speculative Actions reduce latency by 20%. The three constraints ultimately converge into a single Pareto curve, and the overarching trend is moving from humans tuning knobs to models making resource-rational decisions on their own.

aideep-dive

Machine Theory of Mind: How Agents Infer Other Agents' Intentions, Knowledge, and Goals

Inferring another's beliefs/goals/intentions from observed behavior is called Machine Theory of Mind. Three lineages: symbolic BDI, Bayesian inverse planning, and deep learning ToMnet. The biggest controversy in the LLM era is that GPT-4 still trails humans by >10 points on ToMBench — are high scores genuine reasoning or statistical shortcuts?

aiguide

The Three Core Pillars of AI Agents: Context, Cognition, Action

An AI agent is not a black box — it is built from three layers: what it knows (Context), how it thinks (Cognition), and what it can do (Action). Understanding these three layers is the key to grasping why agents are sometimes brilliant and sometimes go off the rails, and how to design a truly effective agent system.

Plan-and-Execute: A RAG Pattern That Plans Before It Acts

For complex queries, have the LLM map out what information is needed and in how many steps — then execute that plan. More systematic than thinking on the fly.