A language model's input is finite, but an agent keeps piling up tool outputs. In week two of ML 2026, Hung-yi Lee splits Context Engineering into three moves: compression (summaries, hard clearing, offloading to files, plus ACON, SUPO, and AgentFold, which make compression smarter), filtering (read only the lines you need, load tools on demand as in MCP-Zero), and finally Agentic Context Engineering, where the LLM decides the next context itself — from Dynamic Cheatsheet and ACE to Recursive Language Models. The most useful idea to take away: a subagent is a form of self-directed compression.
The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.
An agent resends its whole history on every call, so five calls already add up to 80K input tokens; 1,500 OpenHands sessions averaged 78K tokens, 37% of them tool results. Neubig works on two layers: at the model layer, hybrid attention (many local layers, one global) plus length curricula make million-token context possible; at the harness layer, stable prefixes earn cache reads roughly ten times cheaper, and compaction that keeps anchors and externalizes evidence gets past the limit — evaluated by how the agent continues afterwards.
In H1 2026, OpenAI, Anthropic, Letta, and LangChain independently chose Markdown files + indexes over vector databases; write permissions shifted back to humans; forgetting mechanisms appeared but nobody implemented Ebbinghaus; Penfield Labs caught 6.4% wrong answers in LoCoMo; three vendors simultaneously adopted 'Dreaming' for offline memory consolidation. Five trends, one conclusion: memory is not a feature — it is an architecture decision.
Agent memory is not one feature — it is at least four distinct engineering problems: working, episodic, semantic, and procedural. This ten-part series walks through the full design space, from taxonomy to coding agent implementations, platform APIs, open-source frameworks, security attack surfaces, and 2026 trend analysis.
CoALA splits agent memory into working, episodic, semantic, and procedural — but the four-cell taxonomy alone doesn't explain why Claude Code uses Markdown files while Mem0 uses vectors. This post adds six independent design axes (read mode, write timing, fidelity, write authority, forgetting, scope) and a file-to-graph spectrum to map the full design space of agent memory systems in 2026.
All five major cloud platforms shipped agent memory APIs in 2025–2026, but their design philosophies diverge sharply: OpenAI writes memory as files, Anthropic mounts memory as a directory, Google uses vectors with topic classification, AWS combines events with pluggable strategy pipelines, and Microsoft abstracts memory behind context providers. Pricing ranges from free to $0.75/1K records/month; tenant isolation spans from 'your app handles it' to IAM as a first-class citizen.
An AI assistant platform running Opus 4.6 hit stream_stall (90s timeout) twice consecutively when generating docx. Root cause: skill instructions lacked one sentence — 'Write a script' — causing the model to output JS code inline instead of writing a file and executing with node. Claude.ai's official SKILL.md has that sentence, and the model consistently takes the safe path.
A user set retrieval top_k to 15, but the monitoring dashboard showed 5 and the streaming UI flashed 5 before jumping to 15. The same top_k value existed at four layers — LLM tool arguments, runtime, trace DB, and streaming payload — each requiring its own override. Three sequential fixes, each revealing the next layer was also wrong.
An agent with a workspace_browse tool said 'file not found' instead of searching. Anthropic, OpenAI, and Google's official guides all point to the same fix: put trigger conditions and workflows in the tool description. A 2025 study found 97.1% of MCP tool descriptions have quality issues.
Should a sub-agent see the parent's conversation? Fork carries full history but token costs grow exponentially. Fresh saves money but lacks context. Industry consensus: default to Fresh, Fork only when needed, and always pair it with history truncation and result compression.
Traditional requirements scatter across Jira, Slack, and meeting notes, losing fidelity at every handoff. intent.md lets the originator collaborate directly with Claude to produce a human-readable, machine-actionable, version-controlled Markdown proto-spec — from conversation to committed document in hours, not weeks.
Claude Code's plan mode lets engineers produce a reviewable, version-controlled implementation plan (plan.md) before writing a single line of code. Design review shifts from the PR diff to the planning stage, and the cost of course-correcting drops from 'rewriting code' to 'editing a document.'
CLAUDE.md is a context file at the repo root that Claude reads at the start of every session — your team's conventions, commands, architecture patterns, and pitfalls. The course's core advice: if Claude makes the same mistake twice, write it into CLAUDE.md.
Anthropic's Claude Academy offers a free 14-lesson course that takes AI-assisted coding from 'individuals using Claude Code' to 'an organization-wide development lifecycle.' Four core concepts — intent.md, CLAUDE.md, Skills, and Hooks — wire together into a complete AI-native SDLC.
NVIDIA SkillSpector scans agent skills for 71 vulnerability patterns; context-mode sandboxes tool output via MCP to 2% of original size; VoiceStudio runs 16 TTS engines locally with zero cloud dependency; Pydantic AI v2.40.0 adds realtime barge-in and @agent.on_event
Letta extends MemGPT's operating-system analogy but is not a standalone memory API. The runtime persists agent state, editable in-context blocks, conversation history, and external archival memory, while the model can actively curate memory through tools.
Volcano Engine's open-source OpenViking stores agent memory, knowledge, and skills as a viking:// virtual filesystem — browsable with ls, tree, and find. Three-tier loading (L0/L1/L2) averages just 550 tokens per retrieval, boosting LoCoMo memory accuracy from 24–57% to 80–83%.
I turned Anthropic's session-cost advice into a global skill, and the first version made the very mistakes it was meant to prevent: a description stuffed with trigger keywords, hard thresholds based on file counts and minutes, and 'protect the main context' conflated with 'spend fewer tokens overall'. Three rounds later the entry point is one page, details live in references, numeric thresholds became four judgment dimensions, and every claim from a draft post was checked against official docs.
Chroma's controlled study shows that even when it fits, a full context degrades performance. Coding agent vendors have landed on seven different responses: compact, hand off, prune, defer loading, isolate, train it into the model, or change the unit of work. Amp removed /compact outright, Atlassian argues summarization should be a last resort, and Cursor's A/B test measured a 46.9% token reduction. The three real disagreements come down to what each team is measuring.
Most people assume GenAI certifications are built around prompt writing. CCAO-F gives Prompting 14% while Output Evaluation gets 21%; CCDV-F gives Prompt and Context Engineering 11.0%. What actually gets tested is structured output, injection-resistant prompting, dynamic context injection, context compression and caching, prompt lifecycle governance, and proving a prompt change helped — closer to context engineering and software engineering than to writing craft. None of the ten asks you to write a prompt on the spot; they are all multiple choice, so explaining why beats having a feel for it. The single most useful line comes from CCAR-F: when business logic must be guaranteed, 'change the prompt first' is usually the wrong answer.
headroom compresses tool output, logs, and RAG chunks locally before sending them to the LLM, reaching 66K stars in 7 months. agentmemory gives Claude Code, Cursor, Codex CLI and a dozen other coding agents a shared cross-session memory store, hitting 27K stars in half a year. Andrew Ng's team releases OpenWorker, a desktop agent targeting knowledge workers beyond engineers. NVIDIA's labs-OO-Agents reimagines agent abstractions with object-oriented design. Mastra 1.59.0 renames CostGuardProcessor to TokenCostControl (breaking). browser-use 0.13.8 adds first-party OpenClaw skill support.
A VS Code extension open-sourced by Microsoft employees that reads your local Claude Code / Codex / OpenCode session logs. The real payload is 45 Markdown rules: prompts under 30 characters, sending the next message within 15 seconds of receiving 20 lines of AI code, instruction files over 4,000 bytes — turning 'context engineering' into numbers you can argue with.
The course lists four techniques for directing agents: instruction files, hooks, commands, subagents. The instruction file is the only one loaded in full every startup, making it config rather than memory; hooks cover what instructions can't, because a rule can be ignored and a hook cannot; commands are the only one a human triggers. The course also marks just one and a half of seven task steps as human work.
The Agent Skills spec fits in a sentence: a directory containing a SKILL.md. The real design is three levels of progressive disclosure — only name and description load at startup, the body loads on a match, bundled files load on demand. This site's own repo carries 35 skills and 7,893 lines of SKILL.md, and startup still costs only those 35 metadata pairs.
Fall 2026 compresses a full week of prompting into one bullet here and adds RePPIT (Research, Propose, Plan, Implement, Test) and MCP. Two RePPIT rules are worth stealing outright: always ask for exactly two proposals, and never let the instance that wrote the code review it. On the MCP side, Anthropic measured turning tools into code calls dropping 150,000 tokens to 2,000.
Stanford CS146S's Fall 2026 syllabus compresses prompting from a full week into a single bullet, drops the terminal and UI-generation weeks, and adds Agent Skills, Agent-Ready Codebases, Background Agents, and AI-Native Team. Grading moved too: the final project fell from 80% to 50%, with 30% now on open source contributions. This series reads all ten weeks.
Chroma tested 18 frontier models and all of them degrade as input grows — as a cliff, not a slope. Memory failures are usually retrieval failures in disguise. And the real cost of KV cache is bandwidth, not storage: every generated token reads the whole cache.
As tools scale up, selection accuracy doesn't degrade gracefully — it collapses: 4 to 51 tools drops from 43% to 2%, 10 to 100+ drops from 78% to 13.62%. The root fix is to stop stuffing everything in at once — Anthropic's Tool Search Tool uses defer loading plus retrieval to cut 85% of tokens, pushing Opus 4.5 accuracy from 79.5% to 88.1%. Description quality has conditional payoff: negligible in simple scenarios, but correctness jumps from 44% to 50% in multi-tool chaining.
CodeGraph uses tree-sitter to extract a codebase into a local SQLite/FTS5 knowledge graph, letting AI coding agents query the graph instead of scanning files. The official end-to-end benchmark (7 repos, median of 4 runs) averages 35% cost savings and 70% fewer tool calls -- but only if the agent actually walks the graph. Delegating exploration to a file-reading subagent that ignores CodeGraph turns it into pure overhead.
Stop stuffing all your tool descriptions into context at session start. Let the model write code, have the runtime execute it, and let tool definitions enter context only at the import line — Anthropic's GDrive→Salesforce example dropped from ~150K tokens to 2K, and Cloudflare's 2,500-endpoint schema shrank from 1.17M to 1K.
A Skill is a folder with a SKILL.md. Three-layer progressive disclosure lets Claude load details only when needed, eliminating the need to re-explain preferences every conversation.
Agent memory isn't a plugin — it's part of the harness itself. Pick the right memory type, estimate data volume, then decide on the technology. And finally, figure out whether you actually own that memory.
Using my own 30+ RAG/Agent posts to audit the blog itself, I identified a prioritized improvement list spanning content quality, site tech, RAG design fixes, harness infrastructure, and AI agent applications — no phases, just priorities.
There are already 6,400+ .claude/agents/*.md files on GitHub. We dissected 4 representative projects — ChemistryTimes (content production pipeline), claude-sub-agent (document-driven development pipeline), agentic (Temporal.io DAG parallel execution), and vs-copilot-multi-agent (hook-enforced memory persistence) — plus ruflo's enterprise-grade swarm architecture, distilling 6 design patterns and 5 practical trends.
Agent CLIs are not smarter autocomplete tools -- they are AI agents that can read your codebase, execute multi-step tasks, and operate in real environments. Claude Code, Codex CLI, Gemini CLI, OpenCode, Aider, Pi, Kiro, Amp, Cursor CLI... the tools keep multiplying, but they all share a common set of design principles -- understanding these principles is how you actually get good at using them.
AI engineering has gone through three phases: Prompt Engineering (write better instructions) → Context Engineering (feed the right information) → Harness Engineering (design the entire working environment). Each evolution doesn't replace the previous one — it operates at a higher level of abstraction.
Context Engineering is the core concept that replaced Prompt Engineering in 2025: the focus shifted from 'how to ask' to 'what information to provide.' Delivering the right information at the right time into the context window is more effective than upgrading to a stronger model. This post covers the definition, four key strategies, practical techniques, and common failure modes.
AI Agent is not a single technology -- it is an entire architecture system. This article is a systematic navigation: starting from the Agent Three Pillars (Context/Cognition/Action), through the three-stage evolution of AI engineering (Prompt -> Context -> Harness), to eight Multi-Agent design patterns and production-grade Harness infrastructure. Each topic links to a dedicated deep-dive article.
An AI agent is not a black box — it is built from three layers: what it knows (Context), how it thinks (Cognition), and what it can do (Action). Understanding these three layers is the key to grasping why agents are sometimes brilliant and sometimes go off the rails, and how to design a truly effective agent system.