Skip to content

CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop

Sep 30, 20261 min
TL;DRCMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.

🌏 中文版

This post is based on the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is part 21 of the Reading CMU 10-423 series and follows L22 + L26: practical risks and the science of alignment. It covers Lecture 23, "Code Generation / Autonomous Agents," given by Matt Gormley on April 8, 2026.

Official materials used: the schedule, the slides (55 pages), and the inked in-class version (55 pages). The inked version has the same text plus red pen marks from class. The schedule lists no readings for this lecture, and this post does not add any. The course's access grade is A3 (definitions in the global AI/CS course map), but the recordings sit behind a CMU Panopto login, so this post is written entirely from the slides.

The lecture asks a practical question: when a model writes code, or operates a computer on its own, how do you tell whether it did the right thing, and how do you wrap it in a system that can keep trying? The deck has seven parts: applications, evaluation, code models, code-specific techniques, tool calling, coding agents, and autonomous agents. The first four are about the model. The last three are about the system around it.

What code generation is for

The deck opens by contrasting two people pair programming with a person coding alongside an LLM, then lists applications of code models:

  • Autocomplete in IDEs (for example GitHub Copilot)
  • Writing functions or classes from a text prompt or a docstring
  • Reading a large codebase and adding a feature
  • Finding and fixing bugs, writing unit tests
  • Translating code between languages, commenting existing code

On the data side, it gives two examples. Dolma is a 3-trillion-token text dataset assembled from existing sources. The Stack is 3TB of permissively licensed code meant for LLMs. The point: most LLM training data already mixes in many programming languages, so coding ability falls out of ordinary pre-training rather than needing a model trained from scratch.

Evaluating code: from "looks similar" to "actually runs"

Three metrics

MetricWhat it measuresThe deck's verdict
BLEUn-gram overlap with a reference (borrowed from machine translation)Codex paper results show it is not a good surrogate
CodeBLEUA mix of n-gram, syntax tree, and data-flow matchingStill compares against a reference
Functional correctnessHow many unit tests passNow the dominant metric

The shift makes sense: two programs written completely differently can both be correct, and a program one character away from the reference can be completely wrong.

Five benchmarks

BenchmarkWhat it isScale noted on the slides
HumanEvalReleased with Codex, scored by pass@k164 handwritten problems
MBPPPython problems a novice could solve; each has a statement, 3 tests, and a solution974 problems, crowd-sourced
DS-1000Named only, not discussed—
SWE-Bench VerifiedResolve real GitHub issues by producing a patch, scored with the repo's unit tests against the reference fixThe original filtered 90k PRs from 12 popular Python repos down to 2,294 tasks; Verified is a human-validated subset of 500
Terminal-Bench 2.0Interactive command-line tasks, each with a Docker container, English instructions, tests on the final container state, and a reference solution89 tasks, each verified by 3 humans

pass@k is the fraction of problems where at least one of k samples passes all tests. The slide notes that the actual estimate draws more samples to cut variance.

SWE-Bench's filters are worth remembering. A PR has to resolve an issue, has to touch tests, and has to flip at least one test from failing to passing. Together those rules give every task an executable acceptance check. Terminal-Bench 2.0's two example tasks show how varied it is. One asks the agent to build POV-Ray 2.2 from source, install it, and render a test scene to compare against a reference image. The other asks it to read a chessboard from an image and write White's best move to a file in algebraic notation.

Six representative code models

ModelWhat the slides highlight
CodeBERTAn early success: 125M parameters, same architecture as RoBERTa; pre-trained with masked LM plus replaced token detection; example use is natural-language code retrieval
CodexThe original model behind GitHub Copilot: a 12B GPT-3 fine-tuned on 159GB of Python. Starting from pre-trained GPT-3 did not improve final performance, but it converged faster
CodeT5Built on the T5 encoder-decoder and, like T5, trained on many tasks at once
InCoder / FIMAnswers how a left-to-right model can fill in the middle; see below
StarCoderOnce one of the best open code models; pre-trained with FIM, 15.5B parameters, 1 trillion tokens
LongCoderTargets large codebases with sparse attention over long inputs

FIM is the idea in this section most worth understanding. InCoder (April 2022) places a mask token where the code should go. FIM (July 2022) splits code into prefix, middle, and suffix, and trains on:

<PRE> prefix <SUF> suffix <MID> middle

At prediction time you stop at <MID> and the model generates the middle. Because the order is rearranged, the model has already "seen" the suffix before writing the middle, and the ordinary causal mask stays as it is. That is exactly what IDE completion needs: code on both sides of the cursor.

A code-specific trick: unit tests as feedback

The deck first names a counterexample. On ordinary reasoning problems, asking an LLM to check and correct itself usually does not help. Code is different because the model can see unit test output. Feeding failing test messages back and refining at test time works very well.

This observation ties the two halves together. Being able to execute code and get objective feedback is what makes the code domain special, and the coding agents later in the lecture automate exactly that write, run, inspect, revise cycle.

Tool calling: the model asks, the system acts

The slide defines tool calling in five points:

  1. It lets an LLM choose and invoke external functions instead of only generating text.
  2. A tool usually has a name, a description, and a schema for its arguments.
  3. The model decides when a tool is needed, emits a structured call, gets the result back, and keeps reasoning.
  4. It is useful for actions or grounded data, like checking weather, querying a database, sending email, or doing calculations.
  5. The model does not perform the action itself. It asks the surrounding system to run the tool.

Point 5 is the one people skip. Permissions, review, and error handling all live on the system side; the model only makes requests.

The deck then uses Kimi K2 to show how tool skills get trained into a model:

  • Synthetic SFT data: build a large repository of tool specs (real MCP tools plus synthetic ones), then generate agents, tasks, and successful tool-calling trajectories.
  • Joint RL: reinforcement learning in real and synthetic environments, so the model learns tool selection and sequencing from interaction outcomes rather than imitation alone.
  • At inference: every request carries the list of available tools, and the model decides on its own when and how to call them.

Coding agents: a loop that keeps trying

The deck uses a figure from OpenHands to show what a coding agent is and cites CodeAct, where the agent's actions are executable code. The core is a five-step loop:

  1. Understand the goal: read the task, the codebase, and the relevant files.
  2. Plan the next step: pick a concrete action such as inspecting code, editing a file, running tests, or searching docs.
  3. Act: change the code or call a tool.
  4. Observe feedback: read compiler errors, test results, logs, or tool output.
  5. Revise: update its understanding and decide what to try next.

The loop runs until the task is solved, the agent gets stuck, or a stopping condition hits. The slide's compact form: read → plan → edit/run → inspect results → repeat.

What the system prompts look like

One slide reproduces the OpenHands system prompt in full, and the next shows the Codex CLI base instructions. The OpenHands prompt is organized into role, efficiency, file system, code quality, version control, pull requests, and problem-solving workflow. A few concrete rules:

  • If the user asks "why is X happening," answer the question and don't try to fix anything.
  • Every action costs something, so combine commands where possible.
  • Don't create multiple suffixed versions of the same file; edit the original.
  • Don't push to a remote or open a PR unless explicitly asked.

Almost all of these rules deal with the agent doing things the user didn't ask for. It reads like a code of conduct more than a capability spec.

Multiple agents and three hard problems

The multi-agent slide borrows a figure from Anthropic's 2026 Agentic Coding Trends Report, going from a single agent to coordinated teams of agents. Then the deck lists the key challenges in building a coding agent:

ChallengeWhat the slide says
Code searchIf you can't find a bug, you can't fix it; if you don't know where a feature belongs, it's hard to implement
Editing codeThis is the fundamental code generation problem: once the right location is found, generate a patch
Training dataTwo sources: RL rollouts of full solution trajectories, or synthetic trajectories (e.g. SERA); consider the latter if you're low on compute

Autonomous agents: operating the web and GUIs

The last part extends agents from code to web pages and graphical interfaces.

Mind2Web interacts with web pages by reading their HTML directly. It uses two models: a ranking LM narrows the page elements to candidates, and a prediction LLM decides which element to act on and how.

VLM agents look at screenshots instead. According to the slide, these systems usually have two parts: a visual grounding model that decides where to click or type next, and a GUI agent model that steers the whole interaction toward the goal.

Set-of-Mark prompting labels the semantically meaningful regions of an image with numbers, drawn directly into the pixels, so the prompt can refer to them. The slides then pose and answer two questions:

  • How do you find the regions? Use an off-the-shelf segmentation model such as Segment Anything.
  • Does it work with any VLM? Not necessarily. The paper found GPT-4V could interpret and ground itself in the marks, while LLaVA-1.5 and MiniGPT-v2 could not.

SeeClick trains an off-the-shelf VLM (Qwen-VL) to perform individual actions from the screenshot alone.

The final slide covers training agents with RL. There are many ways to complete a task, so imitating one human's attempt is a poor target. RL instead gives a positive reward at the end if the task succeeds.

How the course tests this lecture

  • Quiz 6: in class on April 20 per the schedule, covering L21–L24. The questions are not public.
  • Homework: there are no programming assignments after L15, so nothing maps to this lecture.
  • Practice exam: the Spring 2026 practice exam was released ahead of the March 30 exam. Its 13 sections stop at Scaling Laws and do not cover this lecture.
  • HW623: the paper list for 10-623/723 students includes several papers tied to this lecture, such as ToolLLM, DeepSeek-Coder, WebArena, AutoGen, and CRITIC. How HW623 works is in part 23.

Try this tonight: pick three problems from MBPP, have the model you use most generate 5 solutions for each, run the 3 bundled tests yourself, and compute pass@1 and pass@5. Then paste one failing solution back with its error message and ask for a single fix. You'll see whether the slides' "unit test feedback" claim holds in your hands.

What this post can and cannot confirm

Confirmed: schedule dates and titles, the text of the slides and the inked version, and the papers and URLs the slides cite. Not confirmed: anything said aloud in class (Panopto requires a login), the Quiz 6 questions, and details that appear only in figures (for example the growth chart of LLM-based autonomous agents, the multi-agent architecture diagram, and the SWE-Bench Verified score trend). The slide says "In November 2024, Claude 2.0 only solved 1.96% of issues," but the SWE-Bench paper's arXiv ID dates it to October 2023 (2310.06770). The year looks like a slide typo, so this post keeps the 1.96% figure's source and drops the slide's date.

Further reading: on this site, CME295 Lecture 7: Agentic LLMs goes from RAG and function calling to the agent loop, and CME295 2026 Lecture 6: AI Agents covers context management and harnesses. To read agents as a whole course, see Reading CMU 11-768.

Series navigation: previous L22 + L26: practical risks and the science of alignment | next L24–L26: audio, video generation, and interactive world models | series overview

References