Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.
omp's runLoopBody uses a double while loop: the inner loop drives the core model call → tool execution rhythm; the outer loop, when the agent would stop, drains queued steering / follow-up / asides to decide whether to run another turn. This design solves delivery timing for 'user typing while model streams' and 'background tasks quietly queueing messages'.
omp uses StablePrefix to freeze system prompt + tool specs, AppendOnlyLog for append-only messages, and digest-based longestStablePrefix algorithm. When prune/shake/steering rewrite history, only the tail after the divergence point is resent. This maximizes Anthropic/DeepSeek prompt cache hit rate, fixing the old issue where every turn forced ~40k token re-prefill on llama.cpp (issue #3406).
omp's resolveApproval resolves in three layers: tool declaration → user override → mode tier. Undeclared or malformed approvals default to exec (fail-closed). Tool declares tier + optional policy/override/reason/policyKey; user overrides via tools.approval.<tool>; mode (always-ask/write/yolo) sets auto-allow tier ceiling. Iron laws: tool-side deny and user-side deny can never be crossed by mode; yolo ignores override: true but still honors policy: deny|allow|prompt. bash tokenizes approval: allow must cover whole line, deny/prompt match per segment. Same tool switches read/write via policyKey. checkpoint/rewind are paired sisters. subagent runs headless yolo; parent task is the only auth boundary.
omp doesn't have just one compaction: context-full uses LLM summarization with iterative windows and budget halving retries; snapcompact skips the LLM entirely, rendering history as dense PNG bitmaps for vision models to read — solving no-API-key, low-latency, vision-model-cheaper scenarios; branch summary summarizes the abandoned branch during `/tree` navigation so file ops aren't lost; shake mechanically replaces tool-result text and large fenced/XML blocks with placeholders — an emergency hatch when summary is too heavy and prune isn't enough. Four strategies, distinct failure modes, orchestrated by session maintenance or manual triggers.
OMP's Agent stream does more than print tokens: it separates agent, turn, message, and tool-execution events. Understand the contract to render text deltas, tool progress, and errors without mistaking a partial message for committed conversation state.
This 17-part series takes you from CLI user perspective through pi-mono's Agent Loop, Session Tree, Tool System, Extension System, TUI Architecture, Remote Session, Telemetry, Compaction, and Release process. Ideal for developers wanting to self-host agents, research agent architecture, or contribute to pi.
I'm building my own Python coding agent called looplane. This series dissects the source code of five mature projects — pi, oh-my-pi, opencode, codex, and claude-code — topic by topic, while also comparing them with Looplane's current TUI, external CLI runtimes, local gateway, usage/OTel/session tooling, and Cloudflare slice. Every post follows a fixed five-part structure: design problem → how five projects do it → looplane's choice → academic grounding → improvement roadmap, with evidence cited at file#symbol level.
pi's loop is a double while-loop wrapped in an EventStream; claude-code's source openly says stop_reason is unreliable and uses tool_use blocks observed during streaming as the sole continue signal; codex models a turn as a cancellable SessionTask and records sessions with a dedicated rollout crate. looplane chose an ordering — manifest first, JSONL second — that turns Ctrl-C into verified resumption instead of a rerun. All evidence cited at file#symbol level.
A detailed look at OpenAI's Codex agent loop design: how prompts are constructed, how multi-turn conversations are managed, how prompt caching prevents cost explosions, and how context window auto-compaction works.
The agent loop is a serialized per-session run. The part worth studying is how it handles concurrency: an admitted run records an activeWriterRunId claim, every transcript write supplies expectedWriterRunId, and the commit transaction verifies the match — so a superseded run cannot commit stale data.