Skip to content
All tags

#coding-agent

126 posts

CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop

CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.

CS149 L13: Performance Optimization Beyond the Experts — Halide's Algorithm/Schedule Split, Autoschedulers, and LLM Agents

L13 asks what to do when there are too few people who can write fast code. The slides offer three answers. First, raise the level of abstraction: Halide splits what to compute (the algorithm) from how to compute it (the schedule), so one line of schedule turns the same blur into a tiled, vectorized, multi-core version. Second, intelligent search: because the schedule space is well defined, search plus a learned cost model can generate schedules automatically. Third, the emerging option of LLM agents: have a model write CUDA, run it, read the profiler, reflect, and revise, and let it improve itself with a database of examples or prompt optimization. The last slide leaves you with a question: is the real value in DSL design or in the LLM agent?

Reading NTU ML 2026: HW2, AI Agent as an AI Engineer — an AIDE-Style Tree Search That Lets an Open LLM Build a MyGO & Ave Mujica Face Classifier

HW2 doesn't ask you to write a classifier. You write prompts and a pipeline so that an open LLM running on a Colab T4 (by default a 4-bit GGUF of gemma-3-12b-it) plans, codes, runs, and debugs a 10-class MyGO & Ave Mujica character face classifier on its own. The starter code is adapted from AIDE: an Interpreter runs code, a Node records each version, a Journal forms the solution tree, and the Agent decides whether to draft, debug, or improve next. The first thing worth noticing: the starter's evaluation is empty. Every version is marked metric 1.0 and not buggy, so the tree search picks blindly until you fill it in. The rules are strict: "the LLM agent is your representative", and you may not hand-edit code or prediction files.

CME295 2026 Lecture 6 Preview: AI Agents, from Calling Tools to Managing Context and Tuning the Harness

The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.

Reading CMU 11-768 L6: Coding Agents — From Completing a Line to Fixing a Whole Repo

Neubig's L6 splits coding agents into three layers: train a model that can code (pre-training, mid-training, infilling, RL from test rewards), wrap it in a localize–edit–verify loop with the right editing tools so it can change a repo, then evaluate and train it in SWE-bench-style executable environments. Fixing bugs is only about 15% of a developer's day; the next frontier is tests, CI, and maintenance in the outer loop.

AI Agent GitHub Digest — 2026-09-29

Z.ai's ZCode (7k★) bundles a desktop shell, a browser UI and an agent CLI into one coding-agent workbench. magpie (1.6k★) is a menu-bar app that switches the underlying model for Claude Code, Codex or Gemini CLI in one click. jevgrep (1.3k★) uses semantic search to hand coding agents the right files up front, cutting some of the back-and-forth grep tokens. golive-skill and agent-console round it out with deployment automation and session observability. Claude Code shipped v2.1.284 today, adding Sonnet 5.5 as the default model plus a batch of terminal fixes.

AI Agent GitHub Digest — 2026-09-28

BuilderIO/agent-native (6.9k★) defines an agent's tools and a human's UI as the same code. yynxxxxx/Codex-X (4k★) wraps the Codex CLI in a desktop GUI. career-ops-hq/career-ops (72.9k★) puts an agent to work on job hunting, running entirely inside your local coding CLI, with no website and no résumé uploaded anywhere. No notable framework release today — Pydantic AI v2.51.0 and Claude Code v2.1.283 were both already covered in yesterday's digest.

How Nine Coding Agents Handle Long-Term Memory: From CLAUDE.md to MemFS

Nine coding agents have taken at least four different paths for long-term memory: Claude Code and Codex use Markdown files (agent-written, human-readable), Antigravity CLI inherits Gemini CLI's approval inbox (agent proposes, human decides), and Copilot uses citations with JIT verification (auto-deleted after 28 days unverified). Cursor removed its Memories feature and fell back to human-written Rules. Hermes Agent, OpenClaw, and Letta Code treat memory as a core harness component, not a plugin. Their choices on write timing, forgetting, and cross-team sharing are completely different — and none has published a controlled experiment on whether their memory system actually helps.

Funding Brief|Factory's New Round Hits $200M, Valuation Jumps to $5B

Factory closed a new $200M round at a $5B valuation, more than tripling the $1.5B it carried at Series C five months earlier. This round signals that enterprise buyers now treat autonomous coding agents as a formal line item in the engineering budget, not a demo they're still trialing.

Funding Brief|Cognition Series E $2B+

Cognition closed a $2B+ Series E led by a16z and Accel, pushing its valuation from $26B in May to $48B. This is VCs willing to nearly double their price on the same coding agent company — not a bet on whether AI coding is real, but on whether an 83% run-rate revenue jump in four months can sustain that premium.

Learning from Mature Coding Agents (39): Two Rules for Flicker-Free Tool Result Display in TUIs

Four coding agent TUIs all follow the same two rules: tools maintain constant height (one-line summary, never jumping from 0 to N lines) and never auto-collapse (only user-initiated expand/collapse). looplane violates both.

AI Agent GitHub Digest — 2026-09-05

mattpocock/skills gained 2,757 stars in a single day — the fastest-growing repo on GitHub today. Anthropic's own anthropics/skills and the open-source coding agent anomalyco/opencode are trending alongside it. Meanwhile MCP server reverify ran a benchmark on 71 real Windows system files and found AI has a 97% error rate reverse-engineering binaries from memory — deterministic tools caught every single one. On the framework side, pydantic-ai, agno, and haystack all shipped routine patches today, nothing major.

tech

AI-Native Agent 2026:從 Harness Engineering 到 Skill Engineering 的變革

2026 年 AI Native Agent 時代的核心變化:從 Harness Engineering 到 Skill Engineering,工程師角色從執行者轉向評判者,三個階段演進與五大關鍵技能。

Benchmark Shift|CursorBench: Claude Fable 5.1 Debuts at #1, Bumps Grok 4.6 to Third

CursorBench 3.2: Fable 5.1 Max hits 73.4% (previous leader Grok 4.6 Extra High was 70.8%) and debuts by taking both first and second place; Fable 5.1 Max beats the old leader by 2.6 points while costing only $9.64 per task, 44% cheaper than the prior Fable 5 Max at $17.32; Anthropic's own site still lists only Fable 5, with no official Fable 5.1 announcement

tech

Codex 架構總覽:Rust Monorepo、Bazel 建構、跨平台沙箱

Codex 以 Bazel 管理 140+ Rust crate,核心分為 core/tui/exec-server/protocol 四大塊;沙箱用 codex_sandboxing 統一 macOS Seatbelt、Linux Landlock/bwrap、Windows 沙箱三平台介面;exec-server 以 JSON-RPC + Noise Relay 實現遠端執行。

OMP agent loop: Why two while loops? What the outer 'stopped but woken by steering' layer actually does

omp's runLoopBody uses a double while loop: the inner loop drives the core model call → tool execution rhythm; the outer loop, when the agent would stop, drains queued steering / follow-up / asides to decide whether to run another turn. This design solves delivery timing for 'user typing while model streams' and 'background tasks quietly queueing messages'.

OMP three-layer approval & fail-closed: why undeclared custom tools become exec, and what yolo still blocks

omp's resolveApproval resolves in three layers: tool declaration → user override → mode tier. Undeclared or malformed approvals default to exec (fail-closed). Tool declares tier + optional policy/override/reason/policyKey; user overrides via tools.approval.<tool>; mode (always-ask/write/yolo) sets auto-allow tier ceiling. Iron laws: tool-side deny and user-side deny can never be crossed by mode; yolo ignores override: true but still honors policy: deny|allow|prompt. bash tokenizes approval: allow must cover whole line, deny/prompt match per segment. Same tool switches read/write via policyKey. checkpoint/rewind are paired sisters. subagent runs headless yolo; parent task is the only auth boundary.

OMP hashline edit & noop-loop-guard: Why file edits must be hash-anchored, and how 182/205 byte-identical no-op retries were tamed

OMP replaces traditional line-number patches with hashline: 4-hex content hash + N* syntactic block locators eliminate whitespace drift. noop-loop-guard tracks per-session, per-canonical-path, per-input-hash consecutive no-ops; after 3 (NOOP_HARD_LIMIT) it throws a ToolError so the agent loop sees a tool failure — breaking the model's 182/205 retry loop from issue #2081 where soft hints were completely ignored.

OMP Internals (14): metaharness & Benchmark Infrastructure

OMP's metaharness isn't just a benchmark runner—it's experiment-grade infrastructure for hardware-isolated microVMs, unified auth gateway, and baseline-controlled experimental design.

OMP Internals (8): Provider Quirks & Compat Layer — Why KDL Rules, Not TypeScript

OMP encodes all provider-specific wire behaviors as KDL rules, compiled to JSON and applied by a pure cascade resolver — avoiding TS if-else sprawl, enabling static conflict detection at CI, and making behavior versionable.

OMP streaming internals: the event stream is an agent control plane, not a token stream

OMP's Agent stream does more than print tokens: it separates agent, turn, message, and tool-execution events. Understand the contract to render text deltas, tool progress, and errors without mistaking a partial message for committed conversation state.

OMP Internals Deep-Dive (11): TUI Differential Rendering & Composer

OMP builds a custom TUI engine (packages/tui + packages/coding-agent/src/tui) rather than using React/Ink. Core designs: differential rendering (repaint only changed lines via reference equality), explicit history contract, composer multi-modal input & keybindings system. This article analyzes architecture decisions and implementation details from source code.

OMP vs looplane retrospective: what to borrow, what's over-engineered, what not to touch yet

14 posts in, back to the big picture: append-only context, compaction strategy split, three-layer approval, KDL rule tree, session tree/fork are the five most borrowable; snapcompact, metaharness self-built infra are over-engineered; looplane shouldn't self-build full provider catalog, full TUI, or full collab yet. Philosophy diff: omp = batteries-included in-process, looplane = minimal + external runtime.

pi-mono Deep Dive 13: Agent Harness, Skills, System Prompt Assembly — Building Agent Behavior from Scratch

AgentHarness Core Class, System Prompt Dynamic Assembly Flow, Skills Loading & Formatting, Prompt Templates System, How Harness Decides Tool Availability, Result Handling, Telemetry Schema Registration, Default Harness Construction, Extension Harness Extension.

pi-mono Deep Dive 4: Agent Loop — Double-Loop & Event Flow, From Steering to Follow-up Complete Timeline

Heart of pi-agent-core: agentLoop() → runLoop() double while(true). Inner loop handles tool calls + steering messages; Outer loop handles follow-up + prepareNextTurn (compaction, model switch). Enter = steering (inject after current tool), Alt+Enter = follow-up (inject after agent stops). streamAssistantResponse() partial message updates, tool call parsing, parallel/sequential execution, before/after hooks.

pi-mono Deep Dive 1: pi from a CLI User's Perspective — Install, Modes, Session, Model Switch & Message Interjection

Treat pi as a black box first: 4 run modes, session tree persistence, mid-conversation model switching, Enter vs Alt+Enter message interjection, /tree branch navigation. Builds intuition for the architecture parts that follow.

pi-mono Deep Dive 12: Compaction Deep Dive — Strategy, Token Estimation, Branch Summary, Structured Compaction

Complete Compaction Mechanism: shouldCompact Trigger Conditions (Token Ratio, Message Count), estimateTokens Calculation (Char/Word Approximation), findCutPoint Finding Cut Point (Retain Recent N Turns), generateSummary Generating Summary (LLM Call), prepareCompaction Preparing Context, Branch Summary Generation, Structured Compaction (Extension Custom via fromHook), CompactionEntry Details, fromHook Mechanism, Compaction Settings.

pi-mono Deep Dive 15: Containerization, Sandbox, Permission Model — Gondolin, Docker, OpenShell, Security Boundaries

Why Pi has no built-in permission system, Gondolin Extension (micro-VM), Docker mode, OpenShell policy-controlled sandbox, permission model philosophy, three containerization patterns, security boundary comparison, micro-VM vs container vs process isolation.

pi-mono Deep Dive 7: Extension System — Hooks, Custom Tools, UI Components, Lifecycle Complete Mechanism

Complete Extension system analysis: Extension interface definition, onLoad/onUnload lifecycle, four major Hooks (onAgentStart/onBeforeToolCall/onAfterToolCall/onTurnEnd), five extension points (tools/commands/keybindings/ui/settings), ExtensionRunner load order and dependency resolution, ExtensionAPI capabilities, Dynamic Border, Widget, Dialog, Selector UI components, Extension inter-communication, hot reload mechanism, official example Extensions.

pi-mono Deep Dive 9: Model Catalog, Provider Factory, OAuth & Credential Sync — From Auto-Generation to Cross-Device Sync

pi-ai Model Catalog auto-generation flow, Provider Factory registration with Lazy Loading, OAuth 2.0 + PKCE flow implementation, Credential Store (Keychain/Libsecret/Credential Manager/Encrypted File Fallback), Credential Sync cross-device sync mechanism, Model Scope Diagnostics, ModelResolver parsing logic, CredentialSynchronizationOperation state machine.

pi-mono Deep Dive 2: Monorepo Architecture & Core Abstractions — How 7 Packages Divide Work & Why Dependencies Flow One Way

From user-visible features into architecture: 7 npm packages with clear boundaries, one-way dependency flow, why pi-tui/pi-telemetry have zero deps, how pi-ai encapsulates provider details, lockstep versioning avoiding diamond deps. Builds an 'outside-in' mental model.

pi-mono Deep Dive 3: pi-ai — Unifying 15+ LLM Providers, From Lazy Loading to Auto-Generated Model Catalog

pi-ai is pi-mono's anti-corruption layer: upper layers only see Message/Tool/Context/streamFunction; 15+ providers implement details underneath. This post dissects: unified interface design, Provider Factory Registry, Lazy Loading for tree-shaking, Model Catalog auto-generation, OAuth/API Key unification, Credential Sync, Thinking/Reasoning parameter standardization.

pi-mono Deep Dive 16: Release Pipeline — Lockstep Versioning, Binary Build, Trusted Publishing From Code to npm

Full release flow: Lockstep versioning (all packages same version), CHANGELOG, local smoke, release script, Bun+Node binary build, npm-shrinkwrap, GitHub Actions OIDC trusted publishing, R2 release marker, pi.dev/api/latest-version, announcement verification.

pi-mono Deep Dive 10: Remote Session — Client/Server, JSON-RPC 2.0, WebSocket, Reconnection

pi-protocol JSON-RPC 2.0 Definition, pi-client Connection Management & Exponential Backoff Reconnection, pi-server Session Registry, WebSocket Transport, Heartbeat Mechanism, Session Snapshot, Remote Session Handle, RPC Mode Architecture, Streaming Event Transport, Remote Steering/Follow-up Message Interjection.

pi-mono Deep Dive Series: From Zero to Understanding This Minimal Coding Agent's Complete Architecture

This 17-part series takes you from CLI user perspective through pi-mono's Agent Loop, Session Tree, Tool System, Extension System, TUI Architecture, Remote Session, Telemetry, Compaction, and Release process. Ideal for developers wanting to self-host agents, research agent architecture, or contribute to pi.

pi-mono Deep Dive 5: Session Tree — Append-only JSONL, Branching Without History Mutation, Compaction Logic Full Analysis

SessionManager core: JSONL append-only storage, id/parentId tree formation, branch() moves leaf pointer without mutating history, buildSessionContext() handles compaction entries, createBranchedSession() forks to new file. Complete Entry types: message, thinking_level_change, model_change, compaction, branch_summary, custom, custom_message, label, session_info. Migration v1→v2→v3 details.

pi-mono Deep Dive 11: Telemetry — Vendor-neutral Contracts, Schema Definition, Conformance Tests

pi-telemetry Core: TelemetrySchema Defines Span/Event/Attribute, defineTelemetrySchema Creates TypedSpanStarter, InMemoryTelemetryContext/NOOP_TELEMETRY_CONTEXT Zero-overhead Implementations, Conformance Tests Verify Adapter Correctness, AI/Harness Telemetry Schema Complete Definitions, Attribute Type System, Why Not Use OpenTelemetry Directly.

pi-mono Deep Dive 14: Testing, Quality Gates, Supply-chain Hardening — Faux Provider, Browser Smoke, Biome, tsgo, Shrinkwrap, Trusted Publishing

Testing strategy: Faux Provider (no API key e2e), Vitest unit, Browser Smoke (real browser), Biome lint/format, tsgo type check, Pinned Deps, Shrinkwrap, Install Lock, npm Trusted Publishing, CI pipeline.

pi-mono Deep Dive 6: Tool System — 8 Core Tools, Factory Pattern, Parallel/Sequential Execution, Before/After Hooks Interception

Complete analysis of pi-coding-agent's 8 core tools: ToolDefinition (for LLM) vs AgentTool (execution logic), createToolDefinition/createTool Factory, executionMode determines parallel/sequential, beforeToolCall/afterToolCall interception chain, withFileMutationQueue serializes file writes, truncateHead/Line/Tail output truncation, read/write/edit/bash/grep/find/ls/powershell implementation details.

pi-mono Deep Dive 8: TUI Architecture — Differential Rendering, Component Tree, Layout Engine, CSI 2026 Synchronized Output

Complete pi-tui core analysis: Virtual DOM Diff for flicker-free rendering, Component lifecycle, Layout Engine (Flex-like), CSI 2026 Synchronized Output avoiding partial frame tearing, Keybindings Manager, Alt Screen, Bracketed Paste, Kitty/iTerm2 Image Protocol, built-in components (Markdown, Editor, Selector, Diff, Border, Loader, etc.).

Learning Agent Design from Mature Coding Agents (4): Approval Grading and the Audit Trail

Looplane now grades effects as read/modify/modify_execute/execute and fails closed on unclassified tools. Native MCP tools default to execute unless trusted read-only metadata lowers them. Approval events still land in events.jsonl first and grants can scope to one change set or backend; general command rules and universal sandbox coupling remain unfinished.

Learning Design from Mature Coding Agents (37): Code Mode — Compiling Tool Calls into Batches of Executable Code

looplane now ships a bounded tool-program DSL: read-only programs support list/read/search/diff, repeat, and if_contains; modify/check transactions receive whole-transaction approval and roll back touched paths on failure. This is not arbitrary JavaScript/Python code mode, and transaction execution is not parallel.

Learning Design from Mature Coding Agents (26): Context Compression and Compaction — From Gap to Auditable Baseline

Mature-agent compaction must handle triggers, complete-turn cut points, and recovery. looplane now has an 85% high-watermark, automatic compaction, a deterministic native-loop fallback summary, persisted checkpoints, and workspace-context reinjection. Cross-runtime fallback, model-quality summaries, and live-provider long-session validation remain open.

Learning Agent Design from Mature Coding Agents (27): Cross-Session Memory — From Explicit Remembering to Semantic Recall

omp and claude-code provide cross-session memory while the other references mostly rely on instruction files. looplane now has an explicit remember/list/inject baseline: typed JSONL memories enter prompts across sessions, but retrieval is scope-and-recency only, with no semantic ranking, deduplication, forget command, or automatic extraction.

Learning Design from Mature Coding Agents (28): Dangerous Command Interception and Shell Escalation — Between Allowlists and Always Ask

All five projects combine allow/ask/deny decisions, compound-command inspection, and fail-closed behavior. looplane now has a deny-first classifier, critical floor, shell segmentation, timeout-deny, configured allow/deny rules, and visible policy reasons. Broader syntax coverage and live interactive validation remain open.

Learning Agent Design from Mature Coding Agents: Series Overview — Reading Five Codebases to Build My Own

I'm building my own Python coding agent called looplane. This series dissects the source code of five mature projects — pi, oh-my-pi, opencode, codex, and claude-code — topic by topic, while also comparing them with Looplane's current TUI, external CLI runtimes, local gateway, usage/OTel/session tooling, and Cloudflare slice. Every post follows a fixed five-part structure: design problem → how five projects do it → looplane's choice → academic grounding → improvement roadmap, with evidence cited at file#symbol level.

Hooks, Skills, Plugins: The Three-Layer Extension System of Mature Coding Agents

Hooks govern control flow, skills inject knowledge, and plugins package both. looplane now has opt-in deny-only project hooks, a bounded SKILL.md loader, exact enabled_skills selection, plugin manifests/install/list, and external-runtime projection. Input rewriting, full lifecycle coverage, remote registries, and a mature marketplace remain open.

Learning Design from Mature Coding Agents (36): LSP Integration — Pushing Compiler Diagnostics into Agent Context

looplane can now inject repository diagnostics and open-file state into the next model turn, push typed IDE context over WebSocket, package a VS Code bridge, and supervise long-lived LSP subprocesses through ManagedLspServer. Language-specific initialize/didOpen/didChange adapters and live-editor validation remain open.

Learning Design from Mature Coding Agents (30): MCP Integration — the Standard Socket for Tool Ecosystems

An MCP client must handle transports, tool refresh, approvals, and credential boundaries together. looplane now supports allowlisted stdio, Streamable HTTP/SSE, tools/resources/prompts, tools/list_changed, OAuth metadata/PKCE, and a 0600 credential store. A real authorization-server E2E and MCP-specific confirmation UX remain open.

Learning Design from Mature Coding Agents (35): Model Catalogs and Per-Role Routing — looplane's Role Aliases and Reviewer Lane

looplane now has static ModelRole/ModelRoute candidates, opt-in aliases such as --model @cheap, cross-provider fallback, and a no-tool reviewer lane that runs after verification. Role inheritance/override rules and automatic summarizer, parser, or scout routing remain open.

Learning Design from Mature Coding Agents (6): The ModelProvider Abstraction — Why Wrapping an SDK Is Not Enough

Wrapping an SDK directly buys you three walls within months: usage fields that don't agree, error semantics tied to SDK exception types, and tool-call formats that change per provider. All five reference projects separate 'wire protocol' from 'provider identity' as independent dimensions. Looplane goes further with pydantic canonical contracts (Message/ToolCall/Usage/ModelTurn) plus six protocol adapters, forces the OpenAI SDK's built-in retries to zero, and routes every failure through a classified ProviderErrorKind before any retry policy sees it. Its provider table is deliberately copied from pi's packages/ai — lineage, not coincidence.

Learning from Mature Coding Agents (29): OS-Level Sandboxing

An OS sandbox is the kernel boundary beyond path policy. looplane now ships a fail-closed CommandSandbox: sandbox-exec on macOS, Landlock plus seccomp on Linux, and exit 126 when containment cannot be proven. Coverage still focuses on verification commands, and external CI confirmation remains open.

Prompt Version Control: Changing One Word Can Drop an Eval from 5/5 to 0/5

Looplane's prompt is now `m3-exact-edit-v4`: the version persists into artifacts; core/tool/interaction/runtime/instructions/skills/workspace/memory are composed as stable or dynamic sections; and positive/negative examples cover replace_text, unified diffs, and direct replies. Unit tests pin the structure, while live-eval coverage still needs expansion.

Learning Design from Mature Coding Agents (7): Provider Retry Policy — From One 5xx to Bounded Retry and Fallback

Intermittent NVIDIA NIM 500s exposed Looplane's early gap: classified errors with no retry consumer. SDK retries are now disabled; the harness gives each candidate up to five attempts with jittered exponential backoff and capped Retry-After handling, then can move to an explicitly configured fallback model. Both model.retry and model.fallback enter the event log.

Learning Design from Mature Coding Agents (33): Session Recording and Replay — From Event Logs to Safe Forks

looplane now connects events.jsonl to a deterministic reducer, CLI timeline, canonical JSON, SDK replay, and safe event-point forks. Forking never replays prior tools or model calls; provider/live-runtime validation, redaction, and richer replay hooks remain open.

Learning Design from Mature Coding Agents (32): Subagents and Worktree Isolation — Teaching the Main Loop to Delegate

Mature subagents need roles, bounded fan-out, narrowed permissions, and a result contract. looplane now has native named-role schedules, parallel fan-out, child allowed_paths constrained by the parent, unsafe execution disabled by default, and parent-approved transaction proposals. Persistent background lifecycles, recursion trees, and automatic worktree merging remain open.

Learning from Mature Coding Agents (34): Telemetry and Cost Tracking — You Count Tokens, Then What?

looplane now has CostBreakdown, an explicitly estimated static GPT-5-family price table, per-lane usage/cost, and OTel cost fields. Unknown models still show tokens without invented dollars; broader pricing coverage, authoritative external-CLI bills, and live billing reconciliation remain open.

Learning Design from Mature Coding Agents (18): Toolset Design Philosophy — Drawing the Tool Surface Boundary

Looplane's core surface has grown from seven tools to nine with a read-only `tool_program` and rollback-capable `tool_transaction`; search prefers ripgrep and arbitrary shell remains absent. Native MCP tools join only from allowlisted servers and default to execute approval without trusted read-only metadata.

Learning from Mature Coding Agents (3): Workspace Isolation and Path Policy

Looplane's disposable clone and SafePathPolicy protect the source repo. `--sandbox-checks` can now wrap verification commands with macOS sandbox-exec, Linux bubblewrap, or Landlock, while Cloudflare provides a separate bounded Sandbox slice. Network policy, external-runtime coverage, and production hardening are not yet consistent across those backends.

Learning Design from Mature Coding Agents (38): Agent as a Service — Wrapping Your Loop in Something Other Programs Can Call

looplane now has a Cloudflare Durable Object run resource with async creation, status/cancel/artifacts, live NDJSON, and Last-Event-ID SSE; remote approvals use a separate short-lived capability. Python also provides an attach client and a stateful conversation WebSocket. Production deployment, cross-runtime parity, and full multi-tenant hardening remain unverified.

Looplane remote execution on Cloudflare: Worker, Sandbox, Capability DO, and durable RunSession

Looplane's old synchronous M6 path completed one real deployed coding run. It has since grown into an asynchronous control plane with RunSession, SSE, approvals, cancellation, and artifacts, but that newer path has not been live-revalidated. Audience-separated HMAC capabilities enter the Sandbox while provider credentials stay in the Worker; this is not production-traffic or SLO proof.

Looplane's disposable workspace and run bundle: why the source repository stays untouched

Looplane clones an exact Git commit into a detached-HEAD workspace inside the run directory before a runtime edits or verifies code. The source repository, execution workspace, and run artifacts therefore have distinct boundaries. This provides source isolation and an audit bundle, but it is not an OS sandbox.

Looplane's ExternalCodingRunner: why Codex and Claude Code CLI are external runtimes, not ModelProviders

`ExternalCodingRunner` is Looplane's second runtime lane. The external coding CLI owns its model loop and credentials; Looplane hands off a task and disposable clone, then treats the returned patch as untrusted input and reruns path audit, verification, and the source invariant. This is a capability-bounded handoff, not another `ModelProvider`.

Looplane's ModelProvider multi-gateway: multiple protocols, one canonical contract

Looplane collapses OpenAI-compatible, Responses, Anthropic, Gemini, Workers AI, scripted, and experimental Codex OAuth adapters into one `ModelProvider` contract. The Codex OAuth transport reads SSE but still reduces it inside the adapter into one canonical `ModelTurn`; AgentRunner does not consume token deltas.

Looplane's provider-neutral native loop: from one model turn to a verified terminal state

Looplane's native lane is controlled by AgentRunner: prepare a workspace, request a model turn, execute tool calls, append observations, and enter verification only when the model stops calling tools. Step, wall-time, repetition, token, and cancellation guards can terminate the run independently of the model. Protocol translation belongs to the next article.

Looplane's state-first event journaling: recovering between manifest commits and JSONL appends

Looplane maintains append-only `events.jsonl` and atomically replaced `session.json`, reconciling sequences before crash recovery. The same event contract now supports deterministic replay, canonical JSON replay, fork seeds at a selected sequence, and new workspaces without replaying old side effects; ambiguous `tool.started` or `verification.started` states still hard-fail.

Looplane's tool isolation: path allowlists, strict argv, process groups, and credential-free subprocesses

This article follows one Looplane tool call through its mechanical execution boundary: `SafePathPolicy` for paths and symlink escape, fixed argv with `shell=False`, a sanitized subprocess environment, read-version hashes plus atomic replace for writes, and process-group cleanup at timeout. Permission policy, OS containment, and tool programs are reserved for later articles.

Looplane's TUI and CLI: how a run becomes visible in the terminal

Looplane's TUI and plain CLI are two interfaces over the same runtime paths. The CLI selects a presentation mode from TTY state and flags, runners emit events, and the TUI projects those events into thinking, tool, approval, verification, and terminal states. The screen distinguishes native and external runtimes without treating UI entry points as proof of backend maturity.

A map of Looplane: how one coding-agent task crosses workspaces, runtimes, tools, and events

Looplane turns a coding-agent task into inspectable boundaries: native side effects cross Looplane tools, permissions, and sandboxing, while external runtimes retain their own loops and tools before returning a patch for Looplane audit. This article maps the planned 20-part series.

Looplane context pressure, compaction, and workspace reinjection

Near 85% context pressure, Looplane has two distinct paths: the native loop can apply one bounded deterministic history fallback, while a conversation runtime with native compaction can compact after a completed turn. Both paths re-anchor the next request with workspace context.

Looplane IDE/LSP Context: Diagnostics, Open Files, and the VS Code Bridge

Looplane normalizes up to 200 diagnostics and 32 visible files into bounded, repository-local, untrusted context. Its VS Code and managed-LSP paths supply signals rather than completion, rename, code actions, or full IDE RPC.

Looplane local OS sandboxes: fail-closed execution on macOS, bubblewrap, and Landlock

Looplane can wrap configured local commands and verification in macOS sandbox-exec, Linux bubblewrap, or Landlock/seccomp. A required unavailable backend stops with exit 126 instead of running bare, but external CLIs, MCP/LSP processes, and the entire Looplane process are outside this boundary.

Looplane model roles, fallback, cache hints, and estimated cost

Looplane uses a static model-role catalog and retries or falls back only after retryable provider errors. Cache data is a provider hint plus trace, while cost is a static-table estimate; neither is live routing intelligence or a bill.

Looplane Native MCP: transport, authorization, and approval boundaries

Looplane loads project MCP servers only through an explicit allowlist, projects stdio or Streamable HTTP capabilities into the existing ToolExecutor, and preserves hooks, approvals, timeouts, and cleanup.

Looplane permission layering: how dangerous commands become allow, ask, or deny

Looplane applies a non-bypassable critical floor, evaluates user, organization, and project denies before any allows, and keeps execute operations policy-gated even in dangerous mode. This decides authority; it is not an OS sandbox.

Looplane prompts, instruction precedence, and explicit memory: what the model actually sees

Looplane resolves user and root-to-leaf project instructions before rendering named prompt sections for runtime, skills, workspace state, and the latest 20 explicit memories. The pipeline is traceable and reloadable, but it is not semantic memory and repository text does not become system authority.

Embedding Looplane: SDK, ConversationController, and the WebSocket Boundary

Looplane exposes bounded-run and conversation contracts through a typed 0.x SDK facade. WebSocket attach wraps one prebuilt, controller-owned runtime session rather than providing conversation-ID resume or multi-client routing.

Looplane Skills, Blocking Hooks, and Plugin Packages

Looplane treats skills as bounded repository-local guidance, hooks as opt-in host commands that can only deny, and local plugin manifests as packages for skills and hooks; their authority is deliberately different.

Looplane Subagent Scheduling and Parent-owned Transactions

Looplane normalizes each subagent dispatch into dependency waves of at most four nodes, runs read-only children concurrently in isolated workspaces, then makes the parent repeat hooks, approval, and transaction execution for any modification.

Looplane tool programs, transactions, and safe concurrency

Looplane parallelizes calls only when they are read-only, concurrency-safe, and classified as READ. Tool programs provide bounded read-only repeat and branching, while transactions snapshot and restore possible workspace-file changes; external side effects are not rolled back.

AI Agent GitHub Digest — 2026-08-27

deepseek-ai/deepseek-harness (dsh) uses a Cordis plugin architecture to make models, tools, sandboxes, and memory all swappable components, hitting nearly 200k stars a week after its developer preview launch; PrimeIntellect-ai/prime-agent runs long-lived research coding tasks on a Recursive Language Model architecture, surviving terminal disconnects via a persistent IPython session; liqiwa/mcp-radar automates this very kind of digest by scanning GitHub daily for newly ranked MCP servers. On the framework side, Mastra 1.61.0 adds a crash-resilient background task queue, and ComposioHQ/composio 0.17.0 extends SSRF protection to tool-execution downloads and S3 uploads.

AI Agent GitHub Digest — 2026-08-26

tinyhumansai/openhuman uses a local-first Memory Tree to compress your digital life and orchestrate multiple agents, already at 37k stars in early beta; Vercel Labs' fx is a native coding agent CLI written in Zig at under 8 MiB; NVIDIA open-sources labs-OO-Agents, packing an agent's prompt/tool/workflow into a single Python class; CopilotKit/OpenBot containerizes agents with governance gates — every action is reviewed before execution. Agno v3.0.0 is a major breaking release requiring database migration, and Haystack v3.1.0 adds multi-agent delegation via AgentTool and context compression via CompactionHook.

Learning Agent Design from Mature Coding Agents (2): The Shape of the Agent Loop — Event Streams, Checkpoints, Resume

pi's loop is a double while-loop wrapped in an EventStream; claude-code's source openly says stop_reason is unreliable and uses tool_use blocks observed during streaming as the sole continue signal; codex models a turn as a cancellable SessionTask and records sessions with a dedicated rollout crate. looplane chose an ordering — manifest first, JSONL second — that turns Ctrl-C into verified resumption instead of a rerun. All evidence cited at file#symbol level.

Learning from Mature Coding Agents (13): CLI Ergonomics — Make New Tools Feel Already Familiar

Mature coding-agent CLIs have converged on the same conventions: positional prompt, -p means print, exec is headless, resume is a first-class command, -C changes directory; looplane inherits this vocabulary directly, driving learning cost close to zero.

Learning Design from Mature Coding Agents (10): Edit Tool Trade-offs — unified diff, exact edit, hashline, and whole-file

LLMs break unified diffs on bookkeeping: wrong hunk counts, hallucinated context lines. The five reference projects split into two camps — simplify the diff grammar (Codex drops line numbers), or drop diffs entirely (Claude Code/Pi/OpenCode exact replace); OMP goes further by binding read state into the format via hash anchors. looplane took the minimal-intervention path: keep the guarded apply_patch, add a zero-fuzzy replace_text, and its qwen3:4b eval went from stable failure to 5/5.

Learning Agent Design from Mature Coding Agents (9): External CLIs as a Backend — Where Does the Security Boundary Go?

Every mature coding agent ships a machine interface: codex has `exec --json` plus a full app-server JSON-RPC protocol, claude-code has `-p` with stream-json, and pi/opencode/omp each expose a JSON event stream. Wrapping these CLIs as your backend is the fastest path to subscription-backed coding — but they own their agent loop, their login, and their permission model. looplane's answer: let the external CLI fully own its loop while looplane holds only three things — an isolated working copy, patch audit, and final verification. One runtime never impersonates another.

Learning Design from Mature Coding Agents (22): The Gateway Pattern — Turning Any Provider into an OpenAI-Compatible Endpoint

The ecosystem treats /v1/chat/completions as the lingua franca, but your providers don't all speak it. The five reference projects split into three camps: pi and OpenCode make the client speak every dialect natively so no gateway is needed; OMP builds a real protocol translator (foreign wire → neutral context → provider adapter, no raw passthrough); Codex and Claude Code run proxies that translate nothing and exist purely to force traffic through a controllable path. Looplane copies OMP's boundary but narrows it to one wire in, one out: strictly parse OpenAI Chat into a canonical contract, then dispatch to any ModelProvider — and along the way hit a cross-event-loop client-close bug whose lesson is that provider lifecycles belong to the ASGI lifespan, not the signal handler.

Learning Design from Mature Coding Agents (21): Headless Mode and CI Usage — When Nobody Can Click Approve

The biggest problem when an agent enters CI is approval: no TTY, nobody to click approve. The five reference projects converge on two strategies — delegate permission decisions to the calling program (claude-code's control protocol), or replace approval semantics entirely (codex defaults to Never plus sandboxing, opencode auto-rejects). looplane keeps one AgentRunner loop and injects a different ApprovalPolicy: headless uses HeadlessApprovalPolicy, which never reads stdin so it cannot hang the pipeline, and denies EXECUTE by default — fail closed.

Learning from Mature Coding Agents (14): Onboarding Design — Provider-Aware Init and Instant Verification

A blank config file drives people away; a bad credential discovered too late drives them away faster. All five mature agents treat setup as a first-class state, and looplane adds the step most of them skip: verify the key right after saving it.

Learning Design from Mature Coding Agents (20): The Run Artifacts Contract—What Makes a Run Auditable After It Ends?

After an agent run finishes, 'the model said it's done' is not evidence. Codex splits traces into a manifest + JSONL + payloads bundle, omp mirrors on-disk files into SQLite, pi indexes native session files with runs.jsonl. Looplane picked the strictest option: six fixed files per run, the run is incomplete if any is missing, and patch review reads changes.patch—not anyone's verbal claim.

Learning from Mature Coding Agents (16): Runtime Abstraction and Capability Handshake

Five external CLIs expose five different machine interfaces: JSONL event streams, JSON-RPC handshake, HTTP API, ACP, stream-json. The right way to support them is not one interface that pretends they're identical — it's a narrow runtime boundary plus an honest capability matrix. Availability means installed, not authenticated; protocol drift fails closed.

Learning from Mature Coding Agents (11): Sandboxes and Remote Execution — Deploying on Cloudflare Sandbox

A local sandbox limits the blast radius of an agent on your machine; a cloud sandbox is about moving code safely onto someone else's machine. All five mature projects solve the first problem; only looplane actually deployed the second. Lessons from production: mocks can't catch SSE framing, green CI can't catch a stale wheel, and cleanup paths deserve timeouts just as much as success paths.

Learning Agent Design from Mature Coding Agents (19): Session Persistence and Crash Recovery — Rescuing State After the Agent Dies

All five agents store sessions as append-only JSONL plus some form of single-writer protection, but crash recovery lives in the details: pi repairs torn tails, codex reopens and retries after write failures, and looplane picked a 'manifest first' ordering that reduces the only crash window to one repairable slot. This post dissects each project's write ordering and fail-closed conditions, all cited at file#symbol level.

Learning from Mature Coding Agents (12): Can Small Models Code? — Capability Boundaries and Eval Discipline

Small models don't fail at reasoning first — they fail at format stability: tool-call JSON, diff hunk arithmetic, and context budgets all break. The mature harnesses build evals on real model behavior (pi's model-backed evals, OMP calibrating benchmarks from real session logs, Codex even relaxing its parser for weaker models). looplane picks the narrowest but hardest path: one fixture, five real Ollama runs, a manifest declaring exactly which files and patch fragments count as success — and M2's failure kept verbatim as evidence. Never pass mock off as E2E; never spin partial success into full passes.

Learning Design from Mature Coding Agents (17): Startup Performance and Engineering Discipline — It Was Never the Language

A CLI tool pays its startup cost on every invocation, and performance optimization without a baseline means no regression protection. codex uses daemon reuse and skill snapshot caches; claude-code splits its entrypoint into dynamic imports plus a built-in startup profiler; opencode and omp each maintain lazy-loading discipline; pi does none of it and leans on Bun being fast. looplane is Python — slow by birth — so it applies the full discipline: lazy imports, single-flight disk cache, background controller prewarming, and hyperfine paired benchmarks wired to a CI gate that fails on >10% regression.

Learning Design from Mature Coding Agents (8): The Right Way and the Wrong Way to Use Subscriptions — OAuth and Credential Boundaries

The five reference projects split into three camps on subscription auth. Codex and Claude Code implement OAuth only for their own official clients and store tokens in the OS keyring. pi and OMP directly reuse Claude Code's client ID to implement Pro/Max OAuth — technically feasible, but Anthropic's docs explicitly bar third parties from offering claude.ai login without approval. OpenCode removed its bundled Pro/Max plugins entirely, the cleanest policy precedent in the ecosystem. Looplane's rules: own your grant, never scrape another CLI's credentials, accept third-party OAuth only when the provider clearly supports it, and never copy or forward credentials.

Learning Agent Design from Mature Coding Agents (24): Testing a Moving Agent — fake-CLI Contracts, Recorded Streams, TUI Pilot

An agent's two dependencies — the LLM and external CLIs — are both non-deterministic, but mature projects separate 'the moving parts' from 'the shape of the boundary': codex fakes the Responses API with wiremock plus a scripted SSE server and pins its TUI with insta snapshots; opencode built a VCR-style http-recorder package; pi splits model-backed evals from unit tests into two vitest configs; omp wraps its edit benchmark itself in unit tests. looplane stacks four layers against external CLIs: unit tests, fake-CLI contract tests, recorded-stream integration proofs, and Textual pilot TUI tests. The methodology in one line: record real non-deterministic output, then make deterministic assertions about it.

Learning Design from Mature Coding Agents (15): From Full-Screen TUI to Semantic Transcript

Mature coding agent TUIs never print the event stream directly — they build a typed projection layer first and update it in place. looplane took three steps (full-screen composition, runtime-first dual modes, removing the Ask/Agent split) before two old constraints — non-streaming output and resume-without-replay — were truly lifted.

Learning Agent Design from Mature Coding Agents (5): The Verification Gate — Changed Files Isn't Success, Verified Is

None of the five reference projects enforces 'all declared verification commands pass' at the harness level: pi leaves verification to the model, OpenCode and Codex put it in the system prompt, Claude Code uses a separate adversarial verifier subagent but as a soft contract, and only OMP's cleanse actually runs checks from harness code. looplane takes the hardest path: if files changed, every declared verification command must pass before terminal_reason=verified; with no changes, checks don't rerun (no_changes). Whether to verify is decided by code, not by the model.

Why Python: The Cost and Compensation of Language Choice for Coding Agents

None of the five mature coding agents use Python — pi/opencode/claude-code run on TypeScript, codex rewrote TS into Rust, omp bolted ~80k lines of Rust native crates onto its hot path. looplane still chose Python; the costs are startup performance and packaging, compensated by lazy imports, uv, and Cloudflare Sandbox.

AI Agent GitHub Digest — 2026-08-24

duty1g/x64dbg-mcp-server wraps a reverse engineering debugger as MCP tools, hitting 563 stars in two days; Cripacx/mediagen bakes EU AI Act content marking into an image generation MCP server; QwenLM/qwen-code v0.22.0 publishes full SWE-bench Verified test trajectories with a 77.08% pass rate; open-gitagent/gitagent rewrites its core engine in Rust with agent state living entirely inside a git repo. On the framework side, GitHub's official MCP Server v1.10.0 is a security spring-cleaning — a typo in `--tools` now crashes the server on startup.

Antigravity CLI: Google Replaces a 100K-Star Open-Source Tool with a Closed-Source Go Binary

At Google I/O 2026, Antigravity CLI (agy) replaced Apache 2.0 Gemini CLI with a closed-source Go binary. Technical upgrades — multi-agent orchestration, native sandbox, millisecond startup — but free tier cut 98%, open-to-closed source, 28-day transition window. Community reaction was sharp.

Grok Build: xAI's Rust Coding Agent That Uploaded Your Repo Before Going Open Source

Grok Build is xAI's Rust coding agent — 845K LOC, 8 parallel sub-agents, Arena Mode. May 2026 beta, July open-sourced (Apache 2.0) — but the direct trigger for open-sourcing was a privacy incident: it silently uploaded entire repos (including SSH keys, .env files) to Google Cloud Storage at a 27,800x traffic ratio. The exfiltration code remains in the binary, disabled only by a server-side flag.

Muse Code: Meta's First Coding Agent, Trading Training Rights for a 20x Discount

In August 2026, Meta Superintelligence Labs released Muse Code beta. Closed-source static binary, Muse Spark 1.2 model, parallel persistent sub-agents with worktree isolation. The biggest controversy is pricing: Standard at $1.25/$4.25 per M tokens, or Contributor at $0.10/$0.20 — 20x cheaper, but your code enters Meta's training pipeline.

AI Agent GitHub Digest — 2026-08-23

CopilotKit/OpenBot ships an AG-UI-based 'AI coworker' framework where each agent gets its own computer, hitting 2,289 stars in a week; Bruno's official MCP server (usebruno/bruno-mcp) arrives two months after the community version (Ostico/bruno-mcp-studio); the browser-use team spins off a macOS Harness project that gives LLMs six accessibility primitives to control a Mac directly; opencode, now under Anomaly, has ~199K stars — surpassing Anthropic's Claude Code at ~142K. On the framework side, the MCP TypeScript SDK v2 splits the monolith into 8 sub-packages and follows the protocol's stateless redesign, dropping the session handshake entirely.

aideep-dive

Python Coding Agent M11: Why an Exec Loop Cannot Reproduce the Claude Code Conversation Experience

A Claude Code- or Codex-style TUI depends on long-lived sessions, typed transcripts, and approval at tool boundaries—not a screen full of color.

The H2 2026 Harness War: Eight Frameworks Rewriting, Three Model Makers Entering, 110+ CLIs — How to Make Sense of It

In August 2026, it's not just five frameworks moving. Beyond OMP 2, Pi v2, Opencode 2, dsh, and Claude Code, three model makers — Google (Antigravity CLI), Meta (Muse Code), and xAI (Grok Build) — are building coding agents directly. Add Amp, Cline 2.0, and the Codex CLI Rust rewrite, and eight-plus frameworks are undergoing architecture-level changes simultaneously. Factor in 110+ total CLI tools, and H2 2026 is a divergence period for harness methodology. This article analyzes four architectural approaches, one shared direction, and one emerging trust crisis.

aideep-dive

Runloop: Devbox Infrastructure Built for Coding Agents

Runloop combines isolated microVMs, reproducible images, disk branching, credential proxies, and evals in one coding-agent platform; an official case study reports more than 10,000 concurrent Devboxes in one workload.

DeepSeek Harness (dsh): A Coding Agent Framework That Takes Everything-is-a-Plugin All the Way

DeepSeek Harness (dsh) is DeepSeek's official open-source coding agent framework, released as a v0.1 developer preview on 2026-08-13, accumulating 184,000+ stars in 9 days. Its core is the Cordis plugin kernel — model adapters, tools, agent loop, and UI are all swappable plugins. Four runtime modes, with the ability to use Claude Code and Codex as sub-agents. Web UI first, no native CLI.

OMP 2 (Oh My Pi 2): From Pi Fork to Full Rust Rewrite as an Independent Coding Harness

OMP 2 is no longer a Pi fork. The entire codebase has been rewritten from scratch in Rust, with ~41 crates covering a custom bash engine, GPU-accelerated GUI, embedded CPython 3.14t, gRPC transport, and Kokoro-82M TTS. Currently in pre-release with no stable version yet.

Opencode 2: The Cost of Swapping Bun for Node, Tauri for Electron, and Rebuilding the Entire API

Opencode 2 is a major rewrite led by Anomaly (Dax Raad). Runtime migrated from Bun to Node.js (memory issues), desktop from Tauri to Electron (WebKit perf and Node integration), v1 API intentionally incompatible. New: multi-tab parallel sessions, persistent backend service, HTTP API + SDK. Currently beta, stable estimated ~September 2026. ~200K stars.

Pi v2: AgentHarness API Goes Stable, Earendil Incorporates — Minimalism Enters Its Next Chapter

Pi v0.84.0 (2026-08-06) promotes the AgentHarness v2 API to stable. Lane-based v4 Session model makes operations durable and interruptible. CBOR replaces JSON, Unix sockets replace HTTP. Earendil Inc. (Armin Ronacher's PBC) behind it has secured initial funding. 95.4K stars, still MIT, still minimal.

AI Agent GitHub Digest — 2026-08-19

DeepSeek's open-source agent harness 'dsh' crossed 20K stars within an hour of its 8/13 launch and has since accumulated ~158K stars, with 2000+ plugin proposals flooding in within two days. Its core is a Cordis-powered 'everything is a plugin' architecture that can even call Claude Code and Codex as sub-agents. RightNow-AI reimagines agents at the OS level with Rust (openfang), NetEase Youdao ships a desktop Agent built on OpenClaw (LobsterAI), and PrimeIntellect's prime-agent features a self-improving reasoning loop. CrewAI 1.15.16 adds execution context tracking and flow error logging.

Aider: The Oldest Terminal AI Pair Programmer, and Where Its Maintenance Stands

Aider is a terminal AI pair programmer dating back to 2023 (Python, Apache-2.0, ~48.3k stars), designed against the grain of today's autonomous agents: you control context by hand with /add, every edit becomes its own atomic git commit, and architect/editor mode splits planning from editing across two models. But note the maintenance cadence: the latest PyPI release is 0.86.2 from 2026-02, the last commit was 2026-05, and the site still recommends Claude 3.7 Sonnet and o1.

Amp: The Coding Agent That Defines Itself by What It Deletes

Amp spun out of Sourcegraph in December 2025 as Amp Frontier Corporation, and its npm package moved from @sourcegraph/amp to @ampcode/cli. Its defining trait is deletion: the editor extension, Amp Tab, TODO lists, Fork, custom commands, and public threads have all been removed. Monthly subscriptions only arrived on 2026-07-18 (Megawatt $20, Gigawatt $200); before that it was pay-as-you-go only. The current focus is orbs — remote machines that keep working after you close your laptop.

GitHub Copilot CLI: An Agent That Runs on GitHub the Platform

Copilot CLI went GA on 2026-02-25 and is included in every Copilot plan, Free included. Its differentiator isn't the agent — it's the GitHub integration: a built-in GitHub MCP server that works on issues and PRs, org policies inherited automatically, and an `&` prefix that hands work to the cloud coding agent. Billing runs on GitHub AI Credits (1 credit = $0.01): Pro $10/mo includes $15, Pro+ $39 includes $70, Max $100 includes $200.

omp (Oh My Pi): The Fork That Inverts Pi's Minimalism

omp is a fork of Pi, but it is not just a plugin layer stacked on top: it adds roughly 80,000 lines of Rust, pulling grep, shell, AST, and PTY in-process. Built-in tools go from Pi's 7 to 31, plus 14 LSP ops, 28 DAP ops, and 60+ providers. One codebase, two opposite bets.

Antigravity CLI: How Google Folded Gemini CLI Into a Unified Terminal Agent Harness

Antigravity CLI is a terminal agent Google announced at I/O on May 19, 2026. Written in Go (versus Gemini CLI's Node.js), its binary is called agy, and it shares the same agent harness as the desktop Antigravity 2.0. It is also Gemini CLI's successor — the personal-tier Gemini CLI service ends on June 18, 2026.

aiguide

GitHub Copilot Coding Agent: Assign an Issue to AI and Let It Open the PR

GitHub Copilot Coding Agent lets you assign an Issue to Copilot, which then automatically creates a branch, writes code, runs CI, and opens a PR — all inside a cloud sandbox. The key to success is setting up AGENTS.md; without it, the agent tends to go off track. Best suited for well-defined medium-sized tasks; requires Pro+ (1,500 premium requests/month) or Enterprise plan.

aiguide

Lessons from the Trenches: What AI Native Teams Must Get Right

Not everyone should use a coding agent to modify code directly. AI Native teams need interface specs, test-first development, monorepo, security guardrails, human-in-the-loop, and token budget controls. Building an agent platform layer on top of coding agents and clearly redefining developer roles is the right path forward.

aiproject

Vercel Open Agents: Moving the Coding Agent from Your Laptop to the Cloud

An open-source coding agent reference implementation from Vercel Labs. A three-layer architecture separates the web UI, agent workflow, and sandbox VM — designed as a starting point for teams that want to self-host their own Claude Code or Cursor Background Agent.

Claude Code: A Complete Guide to Anthropic's Terminal AI Coding Agent

Claude Code is Anthropic's agentic coding tool that runs in the terminal, IDEs, Slack, GitHub, and on the web. Its core extension system has six layers: CLAUDE.md (persistent context), Skills (on-demand workflows), Hooks (deterministic automation), Subagents (isolated delegation), MCP (external tool connections), and Agent Teams (multi-agent collaboration).

Codex CLI: A Complete Guide to OpenAI's Open-Source Terminal Coding Agent

Codex CLI is OpenAI's open source terminal coding agent (Rust, Apache-2.0, ~106.6k stars) with MCP, subagents, image input, code review, and Skills. The model line is now GPT-5.6 Sol / Terra / Luna, and the desktop app, CLI, and IDE extension share one config.toml.

Gemini CLI: Once the Most Generous Free Terminal Agent, Now Enterprise-Only

Gemini CLI is Google's open source terminal AI agent (Apache 2.0, ~106.6k stars). It once offered 60 requests per minute and 1,000 per day for free, with a 1M context window. The individual tier stopped serving on 2026/6/18 and Antigravity CLI took over. The project isn't shut down — the repo is still maintained — but it now serves only Gemini Code Assist Standard/Enterprise licenses and paid API keys.

OpenCode: A Complete Guide to the Open-Source AI Terminal Coding Agent

OpenCode is an open-source AI coding agent written in TypeScript (MIT, ~198K GitHub stars, repo at anomalyco/opencode) with a built-in TUI, 75+ LLM providers, LSP integration, a Vim-style editor, SQLite session management, and a desktop app. Free, no subscription, local or cloud models.

Pi Coding Agent: A Minimalist Open-Source Terminal Coding Harness

Pi is a minimalist coding agent by Mario Zechner (TypeScript, MIT, ~93K stars) with just 4 core tools and a very short system prompt — everything else you add yourself via Extensions, Skills, and Prompt Templates. It deliberately omits MCP, sub-agents, plan mode, and permission popups. The repo is now earendil-works/pi and the npm scope is @earendil-works.