Skip to content
All tags

#tool-use

27 posts

CMU 10-423 L23: Code Generation and Autonomous Agents — From pass@k to the Coding Agent Loop

CMU 10-423 Lecture 23 has two halves. The first covers code generation: evaluation moved from BLEU to counting passed unit tests, benchmarks run from HumanEval and MBPP to SWE-Bench Verified and Terminal-Bench 2.0, models run from CodeBERT and Codex to FIM and StarCoder, and the code-specific trick is self-correction driven by unit test output. The second half covers agents: what tool calling is, how Kimi K2 synthesizes tool-use data, the five-step coding agent loop, and web and GUI agents such as Mind2Web, Set-of-Mark, and SeeClick. There is no homework for this lecture; only Quiz 6 tests it.

Reading NCCU Yen-Lung Tsai Generative AI, L09: Why 2025 Was Called the Year of AI Agents — Andrew Ng's Four Design Patterns, with Reflection and Two-Stage CoT Built in AISuite

L09 defines an AI agent in one line: the AI finishes the work you would otherwise do yourself. Yen-Lung Tsai follows Andrew Ng's four design patterns (Reflection, Tool Use, Planning, Multiagent Collaboration) but builds only the two easiest. Demo07a hands a draft between a "writer" and a "reviewer" LLM call. Demo07c splits the Lucky Vicky post generator into "think of five reasons, then write the post", a two-stage CoT. Both use AISuite with Groq and a Gradio front end. LangChain, AutoGen and CrewAI appear only on a further-learning list. The week 9 homework asks you to pick one of the two patterns.

Reading NTU ADL 2025 Fall: Conversational AI and Tool Use — From LU/DST/Policy/NLG to LaMDA, WebGPT, and Toolformer

Dialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

CME295 Lecture 7: Agentic LLMs, or Letting the Model Look Things Up, Call Functions, and Run Its Own Loop

Lecture 7 of CME295 (2025) patches three LLM gaps: RAG fixes knowledge frozen at training time with a two-stage retrieve-then-rerank pipeline; tool calling fixes the inability to act by having a backend execute the function call the model writes; agents chain those calls with ReAct's observe-plan-act loop. The 2026 edition renames it AI Agents and adds context compaction, harness optimization, coding agents, and skills, the biggest rewrite in the course.

CMU 11-768 Lecture 1: An Agent Is a Model in a Loop — the Hard Part Is Making It Work

Lecture 1 of 11-768 strips an agent to its minimum: tool definitions and tool calls are just tokens, the harness parses, executes, and feeds results back into context, and running a ReAct loop makes it an agent. Neubig then lists six capabilities a good agent needs, each of which can be built through training or through the harness, and argues that an agent is a system of harness, sandbox, inference, training, and monitoring — not just a model.

Reading CMU 11-768 L2: How Tool Use Turns Tokens into Actions — Schemas, Constrained Decoding, MCP, and Parallel Calls

Neubig splits tool use into five layers: capabilities, mechanics, constraints, interfaces, and systems. A tool call is just tokens the model emits; the harness parses, validates, and matches results back by call ID. Constrained decoding guarantees form, not correctness. MCP's real value is credential brokering. And the same model served by different providers can swing from roughly 15% tool-call errors to under 0.1%.

aideep-dive

LLM Agent Tool Discovery: Why Agents Don't Use Available Tools, and How to Fix It

An agent with a workspace_browse tool said 'file not found' instead of searching. Anthropic, OpenAI, and Google's official guides all point to the same fix: put trigger conditions and workflows in the tool description. A 2025 study found 97.1% of MCP tool descriptions have quality issues.

Learning Design from Mature Coding Agents (37): Code Mode — Compiling Tool Calls into Batches of Executable Code

looplane now ships a bounded tool-program DSL: read-only programs support list/read/search/diff, repeat, and if_contains; modify/check transactions receive whole-transaction approval and roll back touched paths on failure. This is not arbitrary JavaScript/Python code mode, and transaction execution is not parallel.

Learning from Mature Coding Agents (3): Workspace Isolation and Path Policy

Looplane's disposable clone and SafePathPolicy protect the source repo. `--sandbox-checks` can now wrap verification commands with macOS sandbox-exec, Linux bubblewrap, or Landlock, while Cloudflare provides a separate bounded Sandbox slice. Network policy, external-runtime coverage, and production hardening are not yet consistent across those backends.

Looplane tool programs, transactions, and safe concurrency

Looplane parallelizes calls only when they are read-only, concurrency-safe, and classified as READ. Tool programs provide bounded read-only repeat and branching, while transactions snapshot and restore possible workspace-file changes; external side effects are not rolled back.

aideep-dive

How screenshot-to-code Converts Screenshots to Code: Agent Loop, Asset Extraction, Visual Verification

screenshot-to-code is not a one-shot screenshot-to-HTML tool. Its core is a 30-step Agent Loop with 7 tools — extracting real assets from screenshots, self-verifying with Playwright, and running 4 models in parallel so users pick the best output. 74,500+ GitHub stars, MIT License.

CS224N Lecture 10: Six Components of RAG and Language Agents

Lecture 10 moves from question answering and RAG into language agents, then decomposes them into reasoning and planning, memory, tools, data, and evaluation. An agent is an inspectable loop between a model and external state.

Composio: Who Holds Every User's Token When Your Agent Connects a Hundred SaaS Apps

This site covers MCP thoroughly but has never written about the layer underneath it: when your agent acts for ten thousand end users reading their own Gmail, whose database holds those refresh tokens, who rotates them, who revokes them. Composio is currently the most complete answer — MIT-licensed SDKs, a commercial hosted execution and OAuth layer. It claims 1,000+ toolkits; the managed-auth page actually lists 121 with a Composio OAuth app and 96 that require your own credentials. New pricing effective 2026-08-15: 100K free tool calls, $29/mo Pro. This post takes the authorization model down to an operational level and draws the line between wiring up MCP servers yourself and buying an integration platform.

CS146S Week 1: A Coding Agent Is, Underneath, a While Loop

Week 1 of CS146S is 'build Claude Code in 200 lines' plus a dissection of production system prompts. The agent loop really is that small. The course slides close with four things Claude does underneath, one of them being `<system-reminder>` tags scattered everywhere to stop the model drifting — which appears in no official documentation.

The Protocol Layer: MCP, A2A, ACP, Skills

MCP governs agent-to-tool, A2A governs agent-to-agent, Skills govern reusable knowledge. The test is whether the data changes: if it changes between calls you need MCP; if it's stable enough to write down, a skill file is simpler and has no runtime that can fail on its own.

AI Agent Arxiv Digest — 2026-08-08

Memory Reward Inflation finds that self-improving agents' memory rewards self-inflate — wrong experiences grow more confident over time; LUCID boosts accuracy from 54.0% to 56.9% on BIRD. RoMeRL compresses memory state space with fixed-dimension semantic coordinates, cutting Cold-Q ratio by 80% and LLM calls by 21.1%. ToolLIFT abstracts tool trajectories into function-level workflow graphs, consistently outperforming existing methods on three OOD benchmarks

AI Agent Arxiv Digest — 2026-08-05

ToolLIFT abstracts tool trajectories into function-level workflow graphs, lifting OOD accuracy by 4+ points on average; HyperAgent builds tool-schema hypergraphs with deficit-oriented expansion, beating ReAct by 14.3 points on AppWorld with lower token cost; a multilingual multi-agent planning diagnosis finds that planning grounding failures rise with decreasing language resources, and the TART fix improves scores by 5.6 points on average

aideep-dive

Agent Observability: From OTel Traces to Catching Hallucinations, Tool Misuse, and Infinite Loops

The industry has converged on using OpenTelemetry GenAI semantic conventions to turn every LLM call and tool call into a span. Detecting the three major failure modes then splits into three tracks: faithfulness + semantic entropy for hallucinations, framework-level symbolic guardrails for tool misuse, and max steps + action hash deduplication for infinite loops — all wired into a Final / Trajectory / Single-step three-layer evaluation framework.

aideep-dive

Stop Hand-Tuning Prompts: From GEPA to Tool Descriptions, Automating Agent Behavior Optimization

Automatic prompt optimization (APO) has evolved from APE/OPRO to GEPA: replacing sparse rewards with linguistic reflection, winning over GRPO by ~6pp with 4-35x fewer rollouts. Meanwhile, tool descriptions are the overlooked prompt -- small wording changes can shift tool selection rates by 10x, and Anthropic's experiments show Claude self-rewriting tool descriptions outperforms human experts. These two lines are converging: eval-driven automatic optimization is eating hand-tuned prompts.

aideep-dive

How to Pick the Right Tool from Hundreds: The Collapse Curve of Tool Selection and Engineering Solutions

As tools scale up, selection accuracy doesn't degrade gracefully — it collapses: 4 to 51 tools drops from 43% to 2%, 10 to 100+ drops from 78% to 13.62%. The root fix is to stop stuffing everything in at once — Anthropic's Tool Search Tool uses defer loading plus retrieval to cut 85% of tokens, pushing Opus 4.5 accuracy from 79.5% to 88.1%. Description quality has conditional payoff: negligible in simple scenarios, but correctness jumps from 44% to 50% in multi-tool chaining.

aideep-dive

Auto-Embedding on File Upload Is a Bad Default: A Survey of Adaptive / Agentic RAG and Agentic Parsing

Making 'chunk and embed every uploaded file automatically' the default behavior means making a decision for the LLM that it could have made itself. From Self-RAG (2310.11511) and Adaptive-RAG (2403.14403) to AgenticOCR (2602.24134), the academic trajectory is pushing three layers of decision-making -- whether to retrieve, whether to parse, and how to chunk -- from the ingestion pipeline back to the agent at conversation time.

aideep-dive

Assembling LLM Agent Skills / Tools / Code Interpreter for Real: A Paper Reading Map

The hard part of LLM agents is not building function calling, skills, code interpreter, and document tools individually -- it is assembling them into a system that selects the right tool, writes code when needed, decomposes tasks, verifies results, and resists prompt injection. This post organizes the key papers into six engineering decisions: function calling reliability, tool/skill selection, code-as-action, multi-step planning, skill systems, and safety plus document generation.

aiguide

MCP vs CLI vs API: The Real Boundaries of Agent Tool Interfaces

MCP is not going away, but its effective scope is narrower than most people think. For local development, CLI and raw API almost always beat MCP. MCP's truly irreplaceable niche is the narrow gap of 'cross-agent shared local tool layer.'

aiproject

OpenHarness: A Fully Open-Source Agent Harness Framework

An open-source Agent Harness framework from HKUDS (HKU Data Science Lab) that implements tool calling, skill loading, memory, permissions, and multi-agent collaboration as complete infrastructure, supporting Anthropic / OpenAI / GitHub Copilot API formats.

aiguide

AI Agent Tool Descriptions Shouldn't Be Static: Dynamic prompt() Design Learned from Claude Code

Every one of Claude Code's 45 tools uses a prompt() method that dynamically adjusts based on user type, feature flags, and system capabilities. Applying this pattern to a ReAct Agent, tool descriptions are dynamically generated along three dimensions: orchestrator model capability, locale, and available tools. Small models automatically get few-shot examples; large models save tokens.

OpenClaw's Model Requirements and Provider Ecosystem: Provider, Model, and Runtime Are Three Different Things

OpenClaw's hard requirement for a model is tool use plus a large enough context — onboarding only auto-suggests a local model when it confirms tool support and at least a 16K context window. The easier thing to get wrong is that provider, model, and agent runtime are three separate layers: an `openai/*` ref does not mean Codex.

aiguide

MCP (Model Context Protocol): The Standardized Protocol for AI Agent Tool Invocation

Every AI tool has its own calling format, making integration costly. MCP (Model Context Protocol) is an open standard proposed by Anthropic that unifies the communication protocol between AI Agents and external tools/data sources, enabling tools to be reused across Agents.