Skip to content
All tags

#cmu-11768

14 posts

Reading CMU 11-768 A1: Build an Agent Harness by Hand — One ReAct Loop to Fix Bugs, Compact Context, and Play Chess

CMU 11-768's first assignment starts from an empty ReAct loop: a bash-only CodeAgent fixes a bug in a chess app, context compaction is added to solve a SWE-bench task, and the same loop becomes a ChessAgent that runs a two-ply search with simulate_move, run_python, and a skill. All 100 points are graded by replaying submitted patches and trajectories offline.

Reading CMU 11-768 A2: Writing a Validator for a Data-Visualization Agent — Four Error Families, MCC, and Harbor Verifiers

11-768's Assignment 2 has students use one fixed judge, Qwen3-VL-30B-A3B, to flag four error families in every run of a data-visualization agent, graded by the mean MCC across families on a private set (30% of the assignment). The second half packages students' own tasks as Harbor environments, with one wrong solution the verifier rejects and one that fools it. The theme: the grader you write becomes the RL reward later.

Reading CMU 11-768 AI Agents: An Agents Course You Can't Take Without Having Trained a Language Model, With Three Assignments From Harness to Eval to RL

CMU 11-768 is a new Fall 2026 graduate course on agents taught by Graham Neubig and Daniel Fried. The prerequisite — prior experience training language models — is strictly enforced. Its 23 lectures run from tool calling, context, memory, and planning through SFT, RL, sandboxing, and human-agent interaction. Three individual assignments in the first half build a harness, an evaluation, and a training pipeline; the second half is a team research project. Slides and the first nine lecture videos are public.

CMU 11-768 Lecture 1: An Agent Is a Model in a Loop — the Hard Part Is Making It Work

Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.

Reading CMU 11-768 L2: How Tool Use Turns Tokens into Actions — Schemas, Constrained Decoding, MCP, and Parallel Calls

Neubig splits tool use into five layers: capabilities, mechanics, constraints, interfaces, and systems. A tool call is just tokens the model emits; the harness parses, validates, and matches results back by call ID. Constrained decoding guarantees form, not correctness. MCP's real value is credential brokering. And the same model served by different providers can swing from roughly 15% tool-call errors to under 0.1%.

Reading CMU 11-768 L3: How Long-Context Agents Manage Memory — Hybrid Attention, RoPE Extension, Prompt Caching, and Compaction

An agent resends its whole history on every call, so five calls already add up to 80K input tokens; 1,500 OpenHands sessions averaged 78K tokens, 37% of them tool results. Neubig works on two layers: at the model layer, hybrid attention (many local layers, one global) plus length curricula make million-token context possible; at the harness layer, stable prefixes earn cache reads roughly ten times cheaper, and compaction that keeps anchors and externalizes evidence gets past the limit — evaluated by how the agent continues afterwards.

Reading CMU 11-768 L4: Skills and Memory — How Agents Stop Starting Over

Lecture 4 of CMU 11-768 sorts cross-task experience into episodes, facts, and skills, stored as external artifacts rather than in context or weights. Human-written skills load through SKILL.md and progressive disclosure, lifting the average SkillsBench pass rate from 33.9% to 50.5%. Skills an agent induces itself can be tested before admission when written as code, but break easily on a new website. The hard part is the lifecycle: imperfect judges, over-retrieval, and bloated skill libraries each eat into the gains.

Reading CMU 11-768 L5: Planning — When an Agent Should Think It Through, and When It Should Revise as It Goes

Lecture 5 of CMU 11-768 defines an agent's plan as an explicit, inspectable, revisable representation of intended behavior for this task, and gives four reasons to add planning structure: modularity, environment feedback, long horizons, and control. Fried's own MACU has a manager decompose tasks into a DAG and dispatch parallel sub-agents, raising Odysseys success from 8.5% to 34.0%; on an OSWorld subset, no planning scores 25.0%, an initial DAG with no revisions scores 27.8%, and allowing 10 revisions reaches 58.3%.

Reading CMU 11-768 L6: Coding Agents — From Completing a Line to Fixing a Whole Repo

Neubig's L6 splits coding agents into three layers: train a model that can code (pre-training, mid-training, infilling, RL from test rewards), wrap it in a localize–edit–verify loop with the right editing tools so it can change a repo, then evaluate and train it in SWE-bench-style executable environments. Fixing bugs is only about 15% of a developer's day; the next frontier is tests, CI, and maintenance in the outer loop.

Reading CMU 11-768 L7: How Computer Use Agents See the Screen, Get Graded, and Get Trained

JY Koh breaks computer use agents into three questions: evaluation has moved from single clicks (ScreenSpot-Pro, Mind2Web) to programmatic end-state checks (WebArena, OSWorld), VLM judges, and long-horizon rubrics (Odysseys, OSWorld 2.0); the model is a VLM reading interleaved screenshots and actions; training runs pre-training for grounding → SFT on human and synthetic trajectories → RL in resettable simulated environments.

Reading CMU 11-768 L8: How to Do SFT for Agents — Loss Masks, Trajectory Selection, Data Formats, and the Handoff to RL

Yueqi Song breaks agent SFT into six decisions: compute loss on assistant tokens only (including the stop token); choose trajectories carefully (runs that pass tests can still teach bad habits, and switching teachers or adding new tasks beats sampling more); unify formats with the Agent Data Protocol; watch packing and template drift during training; evaluate in the real harness; and pick the SFT checkpoint for the RL that follows, not for its own best score.

CMU 11-768 Lecture 9: RL Basics — a Policy Gradient Is Just an SFT Loss Times a Weight

Using a guess-a-number-from-1-to-16 game, Daniel Fried frames RL as an extension of SFT: he names SFT's three gaps (task mismatch, no learning from failures, never seeing its own mistakes), then derives ReST, REINFORCE, baselines, and GRPO/DrGRPO in turn. All four compute log-probabilities of the tokens the agent itself generated; they differ only in the weight each token gets.

Reading CMU 11-768 L10: How to Evaluate, Train, and Retrieve for Deep Research Agents

Akari Asai's L10 splits deep research agents into three problems: evaluation has to cover four gaps (search difficulty, domain expertise, long-form answer quality, citation support); training runs mid-training → SFT → RL, with DR Tulu's evolving rubrics as the reward for long-form reports; retrieval should let the retriever see the agent's reasoning, which gets AgentIR-4B to 68% on BrowseComp-Plus with Tongyi-DR.

CMU 11-768 Lecture 11: Advanced RL Algorithms — Credit Assignment, Stable Updates, Reward Hacking, and Distillation

Using a bug-fix coding task, Graham Neubig takes Lecture 9's policy gradient into practice: a critic, GAE, or a PRM to credit individual turns; importance ratios and clipping to handle stale data in async RL; and PPO, GRPO, CISPO, GSPO, and DAPO side by side in one table. The largest share goes to the reward itself — verifier errors, reward hacking, and exploration collapse — before closing with on-policy distillation.