Skip to content
All tags

#test-time-compute

3 posts

CS224R L10: RL for LLM Reasoning and Test-Time Compute

Lecture 10 of CS224R Spring 2026 is a guest lecture by Noam Brown of OpenAI, and it makes one argument: reasoning models open a new scaling dimension by moving compute from training to inference. He starts with his own poker AI work, then uses backgammon, chess, and Go to show that thinking longer at inference time has always paid off. Next comes how LLMs got there: chain of thought, majority voting, o1/o3, GRPO, and DeepSeek-R1-Zero. The second half argues the field needs to rethink itself for large-scale test-time compute: multi-agent systems, evaluation as score versus compute, and the budget assumptions behind safety evaluations. The deck is mostly figures, so this post covers only the points visible on the slides.

Reading NTU ML 2026: Self-Improving AI (Part 1) — AI-Generated Answers, Rewards, and Losses, and How Far Humans Can Step Back

Hung-yi Lee opens his May 8 lecture by admitting that "self-improving AI" has no clear definition: it is a process of humans gradually letting go. He splits machine learning into three steps and checks where the "I" can be replaced by AI. Answers can come from the model's own self-corrections, reward shaping can be written by an LLM, the loss can be set by the model itself (scores, majority vote, entropy), and even the questions can come from a proposer model. But experiments keep showing that with no human at all, progress plateaus or the model trains itself into the ground. A strong AI can already train a weaker one, just not better than humans do. His verdict: in May 2026, AI is "still standing at the bank of the Rubicon."

aideep-dive

Resource Rationality for Agents: Optimal Decisions Across Tokens, Tool Calls, and Latency

Agent decision-making under resource constraints is bounded rationality reborn: Rational Metareasoning uses VOC rewards to save 20-37% of tokens, BATS proves that adding budget without budget awareness is futile, FrugalGPT cascades cut costs by up to 98%, and Speculative Actions reduce latency by 20%. The three constraints ultimately converge into a single Pareto curve, and the overarching trend is moving from humans tuning knobs to models making resource-rational decisions on their own.