Skip to content

AI Agent Interview Prep: From Tool Calling and Memory to MCP and Prompt Caching

Oct 3, 20261 min
TL;DRAgent interviews keep returning to eight topics: the agent loop, structured output, function calling, large tool sets, memory, design trade-offs, MCP versus A2A, and prompt caching. This post strings them into one line (concept, mechanism, how to answer) and corrects when the 2026-07 MCP spec changes actually landed.

🌏 中文版

When an interviewer asks about agents, they rarely want framework names. They want to see whether you can walk the whole chain: how a model says "I want to call a tool", how its output is guaranteed to be machine-readable, what to do when there are too many tools for the prompt, where memory lives, and how to keep cost down. This is part 13 of the "AI Engineer Interview Prep" series, and it ties those questions into one thread: the agent loop first, then Structured Output and Function Calling, then tool scale and memory, and finally MCP, A2A and Prompt Caching.

Every section has the same shape: the concept, then the mechanism or comparison, then a short "How to answer" you can say out loud. The fast-moving parts (MCP spec history, provider caching prices) were checked against primary sources, and the query date is stated.

What an Agent Is: A Loop That Decides Its Own Next Step

In Building effective agents, Anthropic splits agentic systems into two kinds. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths", while agents are "systems where LLMs dynamically direct their own processes and tool usage". The difference is not whether tools exist. It is who picks the next step: hard-coded logic, or the model on the spot.

The "think a step, act a step, read the result, think again" structure comes from ReAct (Yao et al., ICLR 2023). Drawn as a flow, it is the diagram that shows up on every interview whiteboard:

flowchart TD
  Goal["User goal"] --> Reason["Reason: the LLM decides the next step"]
  Reason -->|"emits a tool call"| Act["Act: the runtime checks permissions and runs the tool"]
  Act --> Observe["Observe: the tool result goes back into context"]
  Observe --> Check{"Goal reached?"}
  Check -->|"No"| Reason
  Check -->|"Yes"| Answer["Final answer"]
  Act -.->|"failure or step limit hit"| Stop["Stop conditions: max steps, human approval, honest reporting"]

The point people most often get wrong: the LLM never executes a tool itself. It produces a structured call instruction, and the runtime outside the model does the actual work. As code, the loop looks like this:

messages = [{"role": "user", "content": goal}]
for step in range(MAX_STEPS):                    # max steps: the first safeguard against infinite loops
    reply = llm.chat(messages, tools=tools)
    if not reply.tool_calls:                     # the model asks for no more tools, so it is done
        return reply.text
    for call in reply.tool_calls:
        result = run_tool(call.name, call.arguments)    # the runtime checks permissions before executing
        messages.append(tool_message(call.id, result))  # result goes back into context for the next round
raise StepLimitExceeded

Since 2025 every major vendor has shipped its own agent SDK. For interviews you only need each one's positioning and date:

FrameworkWhenPositioning
OpenAI Agents SDK2025-03Successor to Swarm; core primitives are handoffs, guardrails and tracing
Google ADK2025-04Open-source multi-agent toolkit, Python first
Claude Agent SDK2025-09Renamed from Claude Code SDK; ships subagents and hooks
LangGraphEvolvingGraph state machine with nodes, edges and checkpoints; see the LangGraph guide
CrewAIEvolvingRole-based team model; see the CrewAI guide

More frameworks is not better. The same Anthropic article suggests starting with direct LLM API calls, because frameworks "often create extra layers of abstraction that can obscure the underlying prompts and responses", which makes debugging harder. For a fuller taxonomy, see the AI Agent patterns guide; for the different shapes the loop takes, see agent loop shapes in coding agents.

How to answer

An agent is a system where an LLM decides the next step inside a loop: reason, call a tool, read the result, reason again until the goal is met. The idea traces back to ReAct. The LLM only emits a structured call; an external runtime does the executing, so permissions, step limits and error handling belong in the runtime. If the flow is fixed, a workflow is enough, because an agent trades latency and cost for flexibility.

Structured Output: Blocking Illegal Output at the Token Level

By default an LLM emits free text, but programs need parseable JSON. There are three layers of fix, each more reliable than the last:

ApproachHow it worksLimits
Ask for JSON in the prompt, retry on failureInstructions plus a retry loopNo guarantee; retries add latency and cost
Fine-tune the model to "get used to" the formatTraining exposes it to the target formatProbabilistic; it can still break
Constrained decodingDecoding excludes tokens that are illegal under the schemaNeeds a grammar engine and inference-framework support

The core of constrained decoding is short: use a finite state machine (FSM) or a context-free grammar (CFG) to track where you are inside the JSON, compute the set of legal next tokens, and mask the rest. The Outlines paper describes the FSM-based approach, and Microsoft's guidance is another implementation.

# The core of constrained decoding: at each step, compute the currently legal tokens and set every other logit to -inf
allowed = grammar_state.allowed_tokens()    # derived from the current FSM or CFG state
logits[~allowed] = float("-inf")
next_token = sample(softmax(logits))
grammar_state.advance(next_token)           # advance the state by one step

For example, if the schema requires {"name": string, "age": integer}, then once the output reaches "age": the only legal next tokens begin a number. Quotes and letters get their logits pushed to negative infinity, so the model cannot write them even if it wants to.

The industry standard has long since moved from the old JSON mode to strict JSON Schema. In its 2024-08 announcement, OpenAI published an internal eval: gpt-4o-2024-08-06 with Structured Outputs scored 100% on complex JSON schema following, while gpt-4-0613 scored under 40%. Two caveats apply: it is the vendor's own internal eval, and what is guaranteed is only that the format is valid. Whether the values are correct, or invented, is outside what constrained decoding can see, so you still need semantic validation.

Structured Output and the next section's Function Calling are two sides of one idea: tool arguments are also a structured output that must match a schema.

How to answer

The mechanism is constrained decoding. At every decoding step, the schema and the text so far define the legal tokens through an FSM or CFG, and every other token's logits are set to negative infinity, so the format is guaranteed during generation. That is far more reliable than asking for JSON in the prompt and retrying. One caveat to add: strict mode guarantees valid format, not correct content, so semantic validation is still needed.

Function Calling: The Model States Intent, the Application Executes

Function Calling (also called tool use) is the ability of a model to express "I want to call this function" during a conversation. According to Anthropic's tool use docs, the client-tool round trip works like this: Claude responds with stop_reason: "tool_use" and one or more tool_use blocks, your code runs the operation, and you send the result back in a tool_result block. The same page also distinguishes server tools (such as web search), which run on the provider's infrastructure so you never handle the execution.

The full flow breaks into six steps:

  1. The developer defines the function schema: name, description, and a JSON Schema for the parameters.
  2. The schema travels with the request and becomes part of the context the model sees.
  3. The user asks something, and the model decides whether to answer directly or call a tool.
  4. To call a tool, the model outputs the tool name and schema-conforming arguments instead of natural language.
  5. The application validates the arguments, checks permissions, executes the call, and returns the result as a tool result.
  6. The model reads the result and decides whether to call again or give a final answer.

Why can a model do this at all? Earlier research includes Toolformer (a model that teaches itself when to call APIs) and Gorilla (a model connected to massive numbers of APIs). Commercial models today are trained on this whole behavior: decide whether a tool is needed, pick one, produce schema-conforming arguments, and fold the result into an answer. How each vendor encodes the call internally (special tokens, for instance) is an implementation detail. Saying "the model is trained to emit a call in a specific format" is enough for an interview.

Can it get the call wrong? Yes. The usual failures are picking the wrong tool, fabricating arguments, or not calling at all. The fixes are engineering ones: write tool descriptions with clear trigger conditions (the site has a breakdown in why agents have tools but don't use them), validate arguments against a strict schema, return clear error messages so the model can self-correct, and add human approval for high-risk tools.

How to answer

Function Calling is the model's ability to state which function to call and with what arguments. The developer supplies a schema, the model emits a structured call, the application validates and executes it, and the result goes back to the model. The model itself executes nothing. It can be wrong, so you rely on clear descriptions, schema validation and error feedback, plus human approval for risky operations.

Large Tool Sets: Don't Stuff Them All Into the Prompt

Going from a handful of tools to hundreds hits three problems at once: tool definitions eat the context, the chance of picking the wrong tool rises, and latency and cost grow. Research has put numbers on it: in an MCP stress test, RAG-MCP raised tool selection accuracy from a 13.62% baseline to 43.13% and cut prompt tokens by more than half. The site's how to pick the right tool among hundreds collects more evidence on the collapse curve.

The shared principle is to decouple tool discovery from generation: narrow the set first instead of handing over everything.

StrategyHowCost and limits
Tool retrievalEmbed tool descriptions and take the top-k for each queryRetrieval quality depends on description quality, and the retriever itself degrades at thousands of tools
Deferred loadingTool Search Tool: mark definitions defer_loading: true, then search and expand on demandAdds a search round; whether it preserves caching depends on the provider's implementation
Hierarchy and routingA router picks a domain, then a sub-agent sees only that domain's toolsThe router can be wrong, and it adds a call
Shorter descriptions and namingCompress descriptions, use consistent prefix namespacesDescriptions that are too short make similar tools harder to tell apart
Code ModeThe model writes code that calls tools; definitions enter context only on importNeeds a sandbox; see Code Mode
Multi-agentEach agent owns a tool group, and an orchestrator dispatchesMore coordination and cost; see orchestration patterns

Research offers two more routes. ToolLLM faces more than 16,000 real-world APIs and relies on retrieval to choose among them. ToolkenGPT (NeurIPS 2023 oral) represents each tool as the embedding of a special token, so tools are generated like vocabulary and their descriptions need not sit in the prompt.

One easily missed interaction: tool definitions are part of the cached prefix. Per OpenAI's prompt caching docs, changing tool names, descriptions, schemas or ordering affects the already-cached prefix. So dynamic loading should be designed to append at the end, not rewrite the tool list at the front every round.

How to answer

The core idea is not to hand over every tool at once. Use tool retrieval or deferred loading so only the few relevant tools enter the prompt for a given turn; if the set is messy, route through a hierarchy or split it among specialist sub-agents. Tool descriptions still need to be clear, because both retrieval and selection depend on them. And with Prompt Caching in play, keep the tool list in a stable order so the prefix does not break.

Agent Memory: Short-Term in Context, Long-Term in External Storage

Memory first needs to be split into two layers. Short-term memory is the messages inside the context window, and the problem is that it fills up: common fixes are sliding windows, summary compression and compaction, compared on the site in context compaction design. Long-term memory lives outside the context and, by content, splits into three kinds:

TypeWhat it storesTypical implementationRisk
SemanticFacts and preferencesVector database, knowledge graphStale or contradictory facts
EpisodicPast events and interactionsConversation history store, retrieved by time and relevanceRetrieving irrelevant old events
ProceduralLearned practices and rulesWritten back to the system prompt, a rules file or a skillBad lessons get locked in

Three classic architectures are worth naming. Generative Agents proposed "memory stream + reflection + planning": the agent stores experiences in a stream, then periodically distills them into higher-level reflections. Reflexion has the agent write a verbal reflection after a failed task and feed it into the next attempt's context, with no weight updates. MemoryBank focuses on giving LLMs long-term memory.

What interviewers actually want is not the names but four design decisions: when to write (every turn, at task end, or when the model decides), how to read (inject everything, retrieve by relevance, or let the model query with a tool), how to update and forget, and who can see what (isolation across users). The write gate matters most, because a polluted memory takes effect again in every later conversation:

def maybe_remember(candidate: str, source: str) -> None:
    # Pass a gate before writing: trusted source, no embedded instructions, no conflict with existing memory
    if source == "tool_output":          # tool output is untrusted, so it never goes straight into long-term memory
        return
    if conflicts_with_existing(candidate):
        flag_for_review(candidate)       # on conflict, flag for review instead of overwriting
        return
    memory_store.add(candidate, metadata={"source": source, "ts": now()})

The site has three follow-ups: Agent Memory systems on the evolution from read-only RAG to writable memory, four memory types and six design axes on the design space, and the attack surface of agent memory on why the write gate cannot be skipped.

How to answer

Short-term memory is the context window, kept in bounds with sliding windows or summary compression. Long-term memory sits in external storage and splits by nature into semantic (facts, in a vector store), episodic (past events) and procedural (learned rules). The design has to settle when to write, how to retrieve, and how to update and forget, and it should filter untrusted sources before writing so memory cannot be poisoned. Reflexion and Generative Agents are the two reference architectures people cite most.

Key Considerations When Designing an Agent: Reliability, Security, Cost, Observability

This question has no standard answer; a strong response is categorized and comes with countermeasures. Four dimensions work well:

DimensionTypical problemsCountermeasures
ReliabilityHallucination leading to wrong actions, pretending success after a tool failure, infinite loopsMax steps and loop detection; on failure retry, switch approach or report honestly; human approval for critical operations
SecurityPrompt injection, excessive permissionsLeast privilege, sandboxed execution, treat tool output as untrusted data
Cost and latencyEvery step is one LLM call, and tool results quickly bloat contextUse a workflow when one will do; compress context; Prompt Caching
Observability and evaluationWhen something breaks, you cannot tell which step failedLog every step's reasoning, call and result; build an eval set; separate LLM logic from tool execution logic

Two underlying ideas deserve an extra sentence. First, graduated autonomy: fully automatic for low-risk operations, human confirmation for high-risk ones, and ask when unsure. Second, prompt injection stems from the model flattening instructions and data into one token stream, with no architectural way to tell them apart, so the countermeasure belongs in permissions and trust boundaries and cannot rest on "reminding the model to be careful"; see the single crack behind agent security. For observability, see agent observability and failure detection.

For background reading, there is Lilian Weng's LLM Powered Autonomous Agents and the survey by Xi et al.. On risk assessment, Ruan et al. propose a language-model-emulated sandbox for identifying agent risks without running dangerous tools for real.

How to answer

Organize it into four dimensions: reliability (max steps, honest reporting on failure, human approval for critical operations), security (least privilege, sandboxing, tool output treated as untrusted data), cost and latency (prefer workflows, compress context, use Prompt Caching), and observability and evaluation (log every step, build an eval set). The guiding principle is graduated autonomy: the higher the risk, the more a human steps in.

MCP, Function Calling and A2A: Three Different Layers

These three terms get lumped together, but each covers its own layer. Function Calling is a model-layer capability: how the model expresses "I want to call a tool". MCP (Model Context Protocol) is an application-layer protocol: how an agent discovers and calls tools, and one MCP server can serve any MCP-capable client. A2A is also an application-layer protocol, but it governs how agents discover and collaborate with each other.

flowchart LR
  LLM["LLM<br/>Function Calling: decides which tool to call"] <--> Host["Agent host / MCP client"]
  Host -->|"MCP: tools/list, tools/call"| S1["MCP server: GitHub"]
  Host -->|"MCP"| S2["MCP server: database"]
  Host <-->|"A2A: Agent Card, task"| Remote["Another agent: different framework or vendor"]

They are complementary, not substitutes. In practice the MCP client uses the MCP protocol to fetch the tool list from a server (tools/list), hands the definitions to the model, the model uses Function Calling to decide which one to call, and the client sends the call back through MCP to the server to execute and return a result. The site's protocol layer comparison of MCP, A2A, ACP and Skills goes deeper; its test is whether the data changes between calls.

MCP's timeline is easy to get wrong, so this table follows official sources (checked 2026-10):

WhenEvent
2024-11Anthropic releases MCP
2025-03Spec 2025-03-26: Streamable HTTP replaces the old HTTP+SSE transport, and an OAuth 2.1 authorization framework is added; the same month OpenAI announces Agents SDK support for MCP
2025-11Spec 2025-11-25: clients must implement PKCE (S256) in the authorization flow
2025-12Anthropic donates MCP to the Agentic AI Foundation under the Linux Foundation, co-founded by Anthropic, Block and OpenAI
2026-07Spec 2026-07-28: statelessness, covered below

A common mistake is to credit "Streamable HTTP replaces SSE" and "OAuth 2.1" to the latest spec revision. Both have been the baseline since early 2025. What the 2026-07-28 revision actually changes:

  • The initialize handshake and Mcp-Session-Id are removed, and version and capability information ride along with every request, so the protocol layer no longer keeps a session.
  • New required HTTP headers (Mcp-Method, Mcp-Name) make routing easier.
  • List-type responses can carry caching hints.
  • The old HTTP+SSE transport is formally deprecated.
  • Authorization moves to Client ID Metadata Documents in place of dynamic client registration.

If the protocol layer no longer holds a session, what happens to application state? The official alternative is an explicit handle (such as a basket_id) that the model passes between tool arguments. The same announcement also says the TypeScript and Python SDKs have each passed one billion cumulative downloads.

A2A's timeline is simpler:

Technically, the original design built on HTTP, SSE and JSON-RPC, with an Agent Card (a JSON document) describing an agent's capabilities and endpoint. From v1.0 the data model is separated from the protocol binding, which supports JSON-RPC, gRPC and HTTP/REST, and Agent Cards can be signed.

One last point against over-selling MCP: in local development, a CLI or a direct API call often costs less context than MCP, and what MCP uniquely offers is a tool layer shared across agents. That trade-off is discussed in MCP vs CLI vs API, with an introduction in the MCP primer.

How to answer

Function Calling is a model-layer capability that lets the model state which function to call. MCP is an application-layer standard for how agents discover and call tools, so one server works with any client. A2A governs how agents discover and collaborate with each other. They complement each other: MCP supplies the tool list, and the model uses Function Calling to pick one. Add that MCP's latest direction is statelessness, and that Streamable HTTP and OAuth 2.1 have been the baseline since 2025.

Prompt Caching: An Agent Resends the Same Prefix at Every Step

An agent calls the LLM at every step, and every request carries the system prompt, tool definitions, the full history and earlier tool results. Most of that is identical to the previous step, yet it goes through prefill again.

Prompt caching works by storing the KV tensors for the processed prefix. OpenAI's docs put it plainly: the cache stores KV tensors, not the tokens themselves, and a later request with the same prefix that hits the cache reuses them and only processes what is new. On the research side, Prompt Cache (MLSys 2024) studies modular attention reuse, and PagedAttention (SOSP 2023) underpins KV management on the serving side.

The key rule is that the prefix must match exactly. Anthropic's docs say the cache covers the whole prompt in the order tools, system, messages, up to the breakpoint you mark. A change anywhere in the prefix invalidates everything after it:

flowchart LR
  subgraph hit["Prefix unchanged: later requests hit"]
    direction LR
    A1["tools"] --> B1["system"] --> C1["history"] --> D1["new this round<br/>must be reprocessed"]
  end
  subgraph miss["tools changed midway: everything after the change misses"]
    direction LR
    A2["tools (changed)"] --> B2["system"] --> C2["history"] --> D2["new this round"]
  end

A back-of-the-envelope illustration (simple arithmetic, not a measurement) shows the scale: a 2,000-token prefix, 10 steps, 200 new tokens per step. With no caching, you process 31,000 tokens in total. With every request hitting the cache, newly processed tokens drop to about 4,000, and the rest is billed at the cache-read price, which is not free.

How each provider charges (queried 2026-10; the official pages are authoritative for prices):

ProviderMechanismWriteRead
Anthropiccache_control marks breakpoints, or a top-level field enables automatic caching; 5-minute default lifetime, refreshed on each hit1.25x the base price for a 5-minute TTL; 2x for a 1-hour TTL0.1x; Opus 5.5 is 0.05x, Fable 5.1 and Mythos 5.1 are 0.025x
OpenAIImplicit only up to GPT-5.5; from GPT-5.6, prompt_cache_options.mode selects implicit or explicit, and content blocks mark breakpoints with prompt_cache_breakpoint (at most 4 write points per request, TTL currently 30 minutes only)1.25x from GPT-5.6; no extra write fee on earlier models0.1x from GPT-5.6; varies by model on earlier ones

OpenAI went from charging nothing for writes to charging 1.25x, which explains the value of explicit breakpoints (an inference from the price structure): you avoid paying write fees for a tail that will never be reused.

On the empirical side, Don't Break the Cache evaluated OpenAI, Anthropic and Google across more than 500 agent sessions and found prompt caching can cut cost by 41–80%, with time to first token also falling. The more useful finding is about technique: putting dynamic content at the end of the system prompt and excluding dynamic tool results is more stable than caching the whole thing.

For agent builders, these rules turn directly into actions:

  1. Put stable content first and changing content last: system prompt and tool definitions up front, user data and tool results at the back.
  2. Keep strings that change every request, such as timestamps and request IDs, out of the prefix.
  3. Fix the tool list's order; when loading tools dynamically, append rather than rewrite what comes first.
  4. Watch operations that rewrite earlier history: OpenAI's docs note that compaction replaces earlier conversation content and can invalidate the cache from the first changed token onward.
  5. Read the cache read and write token counts in the API's usage fields to confirm the hit rate instead of guessing.

Also keep the layers straight: Generative Caching and the site's semantic caching cache similar responses at the application layer, which is a different layer from the provider's prefix caching (which caches KV state). They can stack, but they are not the same thing. For a multi-layer cache design aimed at agents, see cache design for a ReAct agent.

How to answer

An agent resends the system prompt, tool definitions and history at every step, so the prefix is highly repetitive. Prompt Caching stores the KV tensors for that prefix, so when the prefix matches, the repeated prefill is skipped, saving both latency and input cost. The practical key is that the prefix must match exactly: stable content first, changing content last, and a fixed tool order. On pricing, writes cost extra and reads are heavily discounted; exact multipliers are on the official pages.

Putting It Together

The eight topics look scattered but form one chain. The agent loop needs the model to emit structured calls a program can parse (Structured Output and Function Calling). Tools and memory make the context keep growing, so you need tool retrieval, layered memory and compression. MCP and A2A standardize "connecting to tools" and "connecting to agents". Prompt Caching then drives down the repeated cost of every loop step. Organizing an interview answer around this chain shows you understand why the pieces relate, which beats reciting them one by one.

Three places tend to cost points in preparation: describing strict mode as guaranteeing correct content, presenting 2025-era MCP changes as 2026 news, and memorizing caching price multipliers without stating when they were checked. All three are mistakes that one look at a primary source avoids.

Questions that keep showing up in public question banks

These questions come from seven public question banks (compared in part 11 of this series, the 12 question banks for AI engineer interviews); only questions that recur across banks and map to a section of this post are kept. "Independent sources" counts overlap between banks, not how often a question comes up in real interviews. The amitshekhar and pallavi banks cite no sources and share 26 near-verbatim questions (likely the same maintainer, so they count as one source), the two KalyanKS banks share an author (also one source), and company tags from those banks are not used here. Only question titles and links to where they appear are listed, with no answers reproduced.

Link labels: om = ombharatiya/AI-Engineer-Interview-Questions, aeg = alexeygrigorev/ai-engineering-field-guide, AIML = alirezadir/AIMLInterviews, amit = amitshekhariitbhu/ai-engineering-interview-questions, pal = pallavi-shekhar/ai-engineering-interview-questions-company-wise, ks = KalyanKS-NLP/LLM-Interview-Questions-and-Answers-Hub.

QuestionIndependent sourcesBank linksWhere it fits in this post
Walk me through the core agent loop. What are the components and stop conditions?4om, aeg, AIML, amitWhat an Agent Is: A Loop That Decides Its Own Next Step
What's the difference between a workflow and an agent?3om, aeg, amitWhat an Agent Is: A Loop That Decides Its Own Next Step
What makes a good tool definition? Give concrete design rules.3om, aeg, amit, palFunction Calling: The Model States Intent, the Application Executes
How do you handle tool failures, retries, and idempotency?3aeg, pal, AIMLKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
You've connected six MCP servers. There are now 130 tool definitions and ~45k tokens of schema in context before the user says a word. What do you do?2om, amit, palLarge Tool Sets: Don't Stuff Them All Into the Prompt
Your agent needs to remember things across sessions. Would you use a vector store or rolling summarisation? Defend the choice.2om, palAgent Memory: Short-Term in Context, Long-Term in External Storage
Explain ReAct (Reasoning + Acting) architecture / prompting.2amit, pal, ksWhat an Agent Is: A Loop That Decides Its Own Next Step
What logic belongs in the orchestrator vs the LLM?2aeg, palWhat an Agent Is: A Loop That Decides Its Own Next Step
What is the Plan-and-Execute agent pattern?2amit, AIMLWhat an Agent Is: A Loop That Decides Its Own Next Step
How does function/tool calling actually work mechanically, end to end?2om, amitFunction Calling: The Model States Intent, the Application Executes
What is MCP, and how does it differ from traditional function calling?2om, amit, palMCP, Function Calling and A2A: Three Different Layers
What are Agent Skills, and when do you package knowledge as a skill rather than a tool, an MCP server, or retrieval?2om, amitMCP, Function Calling and A2A: Three Different Layers
When does multi-agent beat single-agent, and when does it make things worse?2om, amit, palLarge Tool Sets: Don't Stuff Them All Into the Prompt
What types of memory do agentic systems need (working, episodic, semantic, procedural)? How do you design long-term memory without polluting it?2aeg, amitAgent Memory: Short-Term in Context, Long-Term in External Storage
What are the biggest security risks with tool-using agents, and how do you sandbox tool execution safely?2aeg, amitKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
How do you implement human-in-the-loop (HIL) patterns and decide when to trigger human review?2aeg, amit, palKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
Make agent actions reversible, or at least auditable, in production.2om, amit, palKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
Agent orchestration across dozens of SaaS systems: where is authorization enforced and why not in the model?2pal, AIMLKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
How does prompt caching work, and how should it change the way you structure prompts?2om, amit, palPrompt Caching: An Agent Resends the Same Prefix at Every Step
What are the essential components of an agent beyond an LLM?1aegWhat an Agent Is: A Loop That Decides Its Own Next Step
Long-running agent drifts after hours and works on the wrong thing; diagnose and fix.1amit, palKey Considerations When Designing an Agent: Reliability, Security, Cost, Observability
Structured output vs function calling.1amit, palStructured Output: Blocking Illegal Output at the Token Level
How does authorisation work for a remote MCP server, and what do teams get wrong when they implement it?1omMCP, Function Calling and A2A: Three Different Layers
Explain multi-layer caching strategies: retrieval cache, prompt cache, and response cache.1aegPrompt Caching: An Agent Resends the Same Prefix at Every Step

This section lists only question titles and links; go to the original repo for the answers. When banks word the same question differently, the table shows the wording from one of them.

The line-number links for amit, pal and aeg point to the main branch as of 2026-10-03 and can shift after those repos change; if a link lands on a different question, search the original file for the question text.

References

Papers and official docs

Frameworks, protocols and caching

On this site

Question banks (source of questions)