CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop
CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.