Skip to content
All tags

#agent-loop

14 posts

CMU 11-768 Lecture 1: An Agent Is a Model in a Loop — the Hard Part Is Making It Work

Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.

tech

Codex ThreadManager:核心協調者的生命週期、分叉語義與 Subagent 圖譜

ThreadManager 持有 Arc<ThreadManagerState> 統管所有 thread,spawn_thread() 統一處理新建/恢復/分叉/子代理四種啟動路徑;ForkSnapshot 定義 TruncateBeforeNthUserMessage/Interrupted 兩種語義;AgentControl 透過 Weak<ThreadManagerState> 避免循環引用;agent_graph_store 追蹤 ThreadSpawnEdgeStatus::Open/Closed。

tech

Codex Turn 狀態機:TurnContext、StepActivation、Context Manager 與壓縮觸發

TurnContext 在 turn 初始化時捕獲所有設定(模型、審批、token budget),後續 step 透過 StepContext 讀取快照;StepActivation 驗證設定變更不違反 legacy 安全約束;ContextManager 用 Arc<Vec> + 版本號實現 Copy-on-Write 歷史共享;壓縮觸發條件為 token_remaining < threshold,支援 remote v1/v2、local、model fallback 四條路徑。

OMP agent loop: Why two while loops? What the outer 'stopped but woken by steering' layer actually does

omp's runLoopBody uses a double while loop: the inner loop drives the core model call → tool execution rhythm; the outer loop, when the agent would stop, drains queued steering / follow-up / asides to decide whether to run another turn. This design solves delivery timing for 'user typing while model streams' and 'background tasks quietly queueing messages'.

OMP append-only context: Why sync conversation by byte-stable prefix? How Anthropic/DeepSeek KV cache gets protected

omp uses StablePrefix to freeze system prompt + tool specs, AppendOnlyLog for append-only messages, and digest-based longestStablePrefix algorithm. When prune/shake/steering rewrite history, only the tail after the divergence point is resent. This maximizes Anthropic/DeepSeek prompt cache hit rate, fixing the old issue where every turn forced ~40k token re-prefill on llama.cpp (issue #3406).

OMP three-layer approval & fail-closed: why undeclared custom tools become exec, and what yolo still blocks

omp's resolveApproval resolves in three layers: tool declaration → user override → mode tier. Undeclared or malformed approvals default to exec (fail-closed). Tool declares tier + optional policy/override/reason/policyKey; user overrides via tools.approval.<tool>; mode (always-ask/write/yolo) sets auto-allow tier ceiling. Iron laws: tool-side deny and user-side deny can never be crossed by mode; yolo ignores override: true but still honors policy: deny|allow|prompt. bash tokenizes approval: allow must cover whole line, deny/prompt match per segment. Same tool switches read/write via policyKey. checkpoint/rewind are paired sisters. subagent runs headless yolo; parent task is the only auth boundary.

OMP four compaction strategies: context-full / snapcompact / branch summary / shake — what each solves and how they switch

omp doesn't have just one compaction: context-full uses LLM summarization with iterative windows and budget halving retries; snapcompact skips the LLM entirely, rendering history as dense PNG bitmaps for vision models to read — solving no-API-key, low-latency, vision-model-cheaper scenarios; branch summary summarizes the abandoned branch during `/tree` navigation so file ops aren't lost; shake mechanically replaces tool-result text and large fenced/XML blocks with placeholders — an emergency hatch when summary is too heavy and prune isn't enough. Four strategies, distinct failure modes, orchestrated by session maintenance or manual triggers.

OMP streaming internals: the event stream is an agent control plane, not a token stream

OMP's Agent stream does more than print tokens: it separates agent, turn, message, and tool-execution events. Understand the contract to render text deltas, tool progress, and errors without mistaking a partial message for committed conversation state.

pi-mono Deep Dive 4: Agent Loop — Double-Loop & Event Flow, From Steering to Follow-up Complete Timeline

Heart of pi-agent-core: agentLoop() → runLoop() double while(true). Inner loop handles tool calls + steering messages; Outer loop handles follow-up + prepareNextTurn (compaction, model switch). Enter = steering (inject after current tool), Alt+Enter = follow-up (inject after agent stops). streamAssistantResponse() partial message updates, tool call parsing, parallel/sequential execution, before/after hooks.

pi-mono Deep Dive Series: From Zero to Understanding This Minimal Coding Agent's Complete Architecture

This 17-part series takes you from CLI user perspective through pi-mono's Agent Loop, Session Tree, Tool System, Extension System, TUI Architecture, Remote Session, Telemetry, Compaction, and Release process. Ideal for developers wanting to self-host agents, research agent architecture, or contribute to pi.

Learning Agent Design from Mature Coding Agents: Series Overview — Reading Five Codebases to Build My Own

I'm building my own Python coding agent called looplane. This series dissects the source code of five mature projects — pi, oh-my-pi, opencode, codex, and claude-code — topic by topic, while also comparing them with Looplane's current TUI, external CLI runtimes, local gateway, usage/OTel/session tooling, and Cloudflare slice. Every post follows a fixed five-part structure: design problem → how five projects do it → looplane's choice → academic grounding → improvement roadmap, with evidence cited at file#symbol level.

Learning Agent Design from Mature Coding Agents (2): The Shape of the Agent Loop — Event Streams, Checkpoints, Resume

pi's loop is a double while-loop wrapped in an EventStream; claude-code's source openly says stop_reason is unreliable and uses tool_use blocks observed during streaming as the sole continue signal; codex models a turn as a cancellable SessionTask and records sessions with a dedicated rollout crate. looplane chose an ordering — manifest first, JSONL second — that turns Ctrl-C into verified resumption instead of a rerun. All evidence cited at file#symbol level.

aiguide

Inside the Codex Agent Loop: How OpenAI Keeps AI Agents Iterating

A detailed look at OpenAI's Codex agent loop design: how prompts are constructed, how multi-turn conversations are managed, how prompt caching prevents cost explosions, and how context window auto-compaction works.

The OpenClaw Agent Loop: Serialization, Writer Claims, and the Fence That Stops a Stale Turn From Committing

The agent loop is a serialized per-session run. The part worth studying is how it handles concurrency: an admitted run records an activeWriterRunId claim, every transcript write supplies expectedWriterRunId, and the commit transaction verifies the match — so a superseded run cannot commit stale data.