Skip to content

Prompt, Context, Harness: The Three Layers, Their Boundaries, and Evaluation Gates

Oct 3, 20261 min
TL;DRPrompt engineering shapes a single instruction, context engineering decides what the model sees right now, and harness engineering wraps the whole non-deterministic system. Bsharat's 26 principles, Breunig's four context failures, and OpenAI's roughly one-million-line, zero-handwritten-code harness experiment each map to one layer. Evaluation is the piece teams skip most, and it belongs in a CI gate that blocks silent regressions.

🌏 中文版

Interviews for LLM application roles tend to chain the same questions together: why prompt engineering matters, how context engineering differs, what harness engineering is, what happens without evaluation, and how a team should wire LLM development into CI/CD. They look scattered, but they ask one thing: when the model gets it wrong, do you know which layer to fix?

This is part 14 of the "AI Engineer Interview Prep" series. It follows the layers from the inside out: prompts first, then context (which decides what goes into the prompt), then the harness that wraps both, and finally evaluation and CI/CD gates. Each section moves from concept to mechanism or comparison to how to answer in an interview. Where the site already has a full article on a topic, the text links to it.

The Big Picture: Three Layers, Inside Out

The three layers are not successive eras that replace each other. They are three scales. The inner layer asks how to word this instruction. The middle layer asks what the model should see at this step. The outer layer asks how the whole system lets an error-prone model work reliably.

flowchart TB
  subgraph H["Harness: the whole system"]
    direction TB
    subgraph C["Context: everything the model sees right now"]
      P["Prompt: wording and structure of one instruction"]
    end
    T["Tool orchestration, sandbox, permission boundaries"]
    E["Evaluation, tracing, feedback loops"]
  end
  H --> M(("LLM"))

The vocabulary is young. In June 2025, Shopify CEO Tobi Lütke said he preferred "context engineering" over "prompt engineering", and Andrej Karpathy backed the term in his own post. Anthropic's engineering team wrote in Effective context engineering for AI agents that they see context engineering as the natural progression of prompt engineering. The word "harness" spread in February 2026 after an OpenAI article, covered below. For another take on the progression, see the site's From Prompt to Harness: The Three Evolutions of AI Engineering.

LayerCore questionWhat it governsTypical failureFirst fix
PromptHow should this be worded?Instructions, examples, output formatAmbiguity, unstable format, skipped reasoningBe explicit, add examples, ask for step-by-step reasoning
ContextWhat does the model need right now?Retrieval, memory, tool output, historyToo little, too much, or contradictory informationDynamic assembly, compression, isolation
HarnessHow does the system stay reliable?Tools, sandbox, constraints, evaluation, observabilityOverreach, infinite loops, silent regressionsDeterministic checks, feedback loops, gates

The most practical use of the boundary is as a diagnostic order. When output is wrong, first check whether the wording is ambiguous, then whether the model received the information it needed, and only then whether the system guards against errors. Outer-layer fixes cost more, but they remove a whole class of failures at once.

How to answer in an interview

Start with containment in one sentence: a prompt is what you do inside the context window, context decides what enters the window, and the harness packages context, tools, permissions, and evaluation into a reliable system. Add that this is a change of scale, not a replacement, so earlier techniques still apply. If pressed, close with "diagnose from the inside out"; that ordering is usually what the interviewer wants to hear.

Layer 1, Prompt: Make One Interaction Clear

Prompt engineering is the practice of designing, refining, and iterating on inputs to steer a model toward the output you want. Liu et al.'s survey frames it as a paradigm shift in NLP: from "pre-train then fine-tune" to "pre-train, prompt, predict," where behavior changes without retraining. The Prompt Report catalogs 58 text-based prompting techniques and works as a map of the field.

Six factors usually come up when interviewers ask what affects output quality:

FactorWhat it doesEvidence and limits
Instruction clarityLeaves less for the model to guessBsharat et al. proposed 26 principles; on their own ATLAS benchmark, GPT-4 response quality rose 57.7% on average (human-rated, GPT-4 and that benchmark only)
ContextNarrows the search space and supplies facts the model never trained onLewis et al.'s RAG injects retrieved knowledge into the prompt, the canonical example
RoleAdjusts tone, expertise level, and viewpointZheng et al. tested 162 personas; adding a persona gave no consistent gain on factual questions
Few-shot examplesShows the input-to-output mappingThe GPT-3 paper showed new tasks can be done from a few examples without fine-tuning
Output formatImproves parseability, cuts post-processingGuaranteeing format takes a constraint mechanism such as structured output, not just asking nicely in the prompt
Reasoning guidanceGets the model to write intermediate stepsIn Wei et al., chain-of-thought on PaLM 540B lifted GSM8K from 17.9% to 56.9%

Three points on this table are easy to overstate.

First, 57.7% is response quality, not accuracy, and it is human-rated, limited to GPT-4 and the ATLAS benchmark. Carry those qualifiers when you cite it, or you are inflating the paper's claim.

Second, role prompts are often treated as a shortcut to correctness. The Zheng et al. paper even reversed its conclusion between versions: the first version said interpersonal roles helped, while the latest version says adding a persona did not improve performance and the effect looks close to random. Roles suit tone and style. They do not replace supplying data.

Third, separate "valid" from "correct" when specifying format. OpenAI's Structured Outputs, introduced in August 2024, makes output follow a developer-supplied JSON Schema. That solves format. Whether the content is right still needs separate verification.

Chain-of-thought also has a zero-example form: Kojima et al. found that appending "Let's think step by step" beats standard zero-shot prompting on several reasoning tasks.

A solid prompt roughly contains role, background, task and constraints, examples, output format, and reasoning guidance, though not every prompt needs all six. Anthropic's advice runs the other way: start with a minimal prompt and add instructions and examples only for failures you actually observe. For the iteration workflow, read the site's Prompt Engineering in Practice. For versioning prompt changes against evals, see Prompt Versioning: One Word Can Drop an Eval from 5/5 to 0/5. Prompt design for RAG is covered in RAG Prompt Engineering.

How to answer in an interview

Give the definition and the paradigm shift, then pick the three factors that matter most: instruction clarity, supplying data, and examples plus reasoning guidance. Volunteering the limits of role prompts earns more credit than reciting the list of six. Finish with how to improve prompts systematically: build a test set, change one thing at a time, run the evaluation, and check for regressions instead of going by feel.

Layer 2, Context: Decide What Enters the Window

Definition and boundary

Context engineering aims to put the right information and tools in front of the model at the right time and in the right format. Karpathy's phrasing is filling the context window with just the right information for the next step, which he breaks into task descriptions, few-shot examples, RAG, tools, state and history, and compaction, with both science and art involved. Anthropic's definition is more engineering-flavored: curating and maintaining the optimal set of tokens during inference, including everything that reaches the model beyond the prompt.

It differs from prompting in three ways. In scope, prompting asks how to word something while context asks what the model needs access to. In time, prompting optimizes one interaction while context reasons about sequences: what earlier turns left behind and which tool outputs should still be there three steps later. In maturity, context engineering reads more like system design. Karpathy's LLM-as-operating-system analogy was packaged by LangChain as "the LLM is the CPU and the context window is the RAM." That sentence is LangChain's phrasing, not a quote from Karpathy's post, so attribute it carefully.

On the academic side, Mei et al.'s survey analyzes more than 1,400 papers and proposes a taxonomy, which makes it the best scholarly citation available. Two 2026 preprints, Calboreanu's practitioner methodology and Vishnyakova's enterprise multi-agent architecture, are single-author. The first is an observational study of 200 interactions with no control group, so its evidence is limited and it works only as background. On the industry side, Gartner's March 2026 data and analytics predictions release lists "the need for context" among the areas AI will affect.

What context is made of

At any moment, an agent's context usually holds these parts:

PartContentCommon problem
System promptRole, rules, boundariesToo rigid or too vague
User inputThis turn's request, from a person or an upstream agentRequirements added turn by turn that contradict each other
Conversation historyEarlier exchangesGrows without bound
Retrieved knowledgeSnippets from a vector store, search, or APIsRelevant but unusable, or ranked wrong
Tool descriptionsAvailable actions and parameter schemasToo many tools, overlapping descriptions
Task metadataUser attributes, permissions, constraintsMissing data leads to overreach or off-target answers
ExamplesFew-shot input and output pairsStuffed with edge cases
Long-term memoryPreferences and conclusions kept across sessionsStale or poisoned

LangChain also offers a coarser three-way split into instructions, knowledge, and tool feedback, which is easy to remember. The site's Context Engineering: Why Your AI Agent's Problem Is Information, Not the Model has the full diagrams and examples.

How poor design breaks things

The intuitive split has three cases. Too little information forces the model to guess, which produces hallucination. Too much dilutes attention and raises cost and latency. Contradictory information leaves the model unsure whom to follow. Drew Breunig splits the "too much and conflicting" cases into four failures:

FailureDescriptionA concrete example
PoisoningA hallucination or error enters the context and is referenced repeatedlyThe Gemini 2.5 technical report describes a Pokémon-playing agent whose goal fields were polluted with wrong game state, and the error took a long time to undo
DistractionThe context grows so long the model leans on its historyThe same report saw the agent favor repeating past actions once context went well beyond 100k tokens
ConfusionSuperfluous content gets used in the answerOn the 46-tool GeoEngine benchmark, a quantized small model failed even though the context was within its window
ClashNew information contradicts earlier informationA Microsoft and Salesforce multi-turn study split full instructions across turns, and average performance dropped 39%

Position matters too. Lost in the Middle (TACL 2024) found that performance is highest when relevant information sits at the start or end of the input and degrades noticeably when it sits in the middle of a long context, even for models built for long inputs. Anthropic cites Chroma's context rot work for the same phenomenon: as tokens accumulate, the model's ability to accurately recall from them declines, so context should be treated as a finite resource with diminishing returns.

That answers the common follow-up, "if windows keep growing, do we still need context engineering?" Yes. A bigger window lets you fit more in. It does not mean the model uses it well, and longer inputs cost more and run slower.

How architecture raises context quality

The guiding principle is Anthropic's: find the smallest set of high-signal tokens that maximizes the chance of the outcome you want. In architecture, that comes down to five moves.

flowchart LR
  Q["Request arrives"] --> S["Select: retrieval, memory, tools"]
  S --> K["Compress and order: rerank, summarize, key facts at the edges"]
  K --> A["Assemble context"]
  A --> L["LLM inference"]
  L -->|"Tool results and new findings"| W["Write out: notes, memory, state"]
  W --> S
  1. Dynamic assembly. Context is not a static template. It is the output of code that runs before the main LLM call and decides, per task, what to include.
  2. Retrieval, reranking, compression. Fetch candidates, use a reranker to keep the few most useful snippets, and summarize into key points when needed. The site's RAG Evaluation Frameworks and Tool Selection shows how to measure this stage.
  3. On-demand retrieval. In the just-in-time approach Anthropic describes, an agent holds lightweight identifiers such as file paths or queries and loads data through tools only when needed. The cost is speed compared with precomputed retrieval, and the design has to teach the model how to look.
  4. Compression and pruning. Long tasks can use compaction (summarize, then restart the window), structured notes, and subagents. Anthropic's compaction keeps architectural decisions and unresolved issues and drops redundant tool output. The hard part is choosing what to keep, because over-aggressive compaction loses details whose importance only shows up later.
  5. Isolation. In LangChain's write, select, compress, isolate framework, isolation means giving multiple agents clean windows or keeping large objects in a sandbox and returning only a summary.

Also treat observability as a design requirement: log the prompts actually sent, the snippets retrieved, and the outputs. Otherwise you cannot tell which part failed.

How to answer in an interview

Define it as "the right information and tools, at the right time, in the right format," and distinguish it from prompting as "wording versus information, single turn versus sequence." List six to eight components, split failures into too little, too much, and conflicting, and add Breunig's four modes and Lost in the Middle for depth. Pick three architectural strategies and go deep: dynamic assembly, compression and pruning, and isolation. Close by saying token cost and latency are design variables.

Layer 3, Harness: Engineering Around a Non-Deterministic Core

Definition and role

The common definition of a harness is everything in an AI agent except the model. LangChain's formulation is Agent = Model + Harness, and Böckeler's analysis on Martin Fowler's site cites the same equation. The relationship to the other layers is containment: a prompt is a technique a harness may use, context management is one of its duties, and the harness also owns tool execution, permission boundaries, error handling, and the whole interaction loop.

What separates this work from traditional software engineering is the non-deterministic core. A harness must expect the model to say or do unexpected things and handle them gracefully. Karpathy's post already hints at this: context engineering is one small part of a "thick layer" of software that also includes control flow decomposition, model dispatch, guardrails, security, evals, and parallelism, which is nearly the harness checklist.

OpenAI's case and Böckeler's grouping

In February 2026, OpenAI's Ryan Lopopolo published Harness engineering: leveraging Codex in an agent-first world. By OpenAI's own account (not independently audited), a team used Codex to build an internal product of roughly one million lines of code in about five months, with around 1,500 pull requests and zero hand-written lines. The article describes the engineer's job as designing environments, specifying intent, and building feedback loops. The site has a full walkthrough: OpenAI Wrote a Million Lines with Codex: Harness Engineering in Practice.

Birgitta Böckeler's first-thoughts memo on Fowler's site grouped the OpenAI team's harness into three categories. She notes in the text that this grouping is her interpretation, not headings from OpenAI's article:

  • Context engineering: a continuously enriched in-repo knowledge base, plus agent access to dynamic information such as observability data and browser navigation.
  • Architectural constraints: enforced not only by LLM-based agents but also by deterministic custom linters and structural tests.
  • Garbage collection: agents that run periodically to find documentation inconsistencies or violations of architectural constraints, fighting entropy and decay.

That memo was later superseded by her full article, which switches to two axes: guides (feedforward controls applied before the agent acts) and sensors (feedback controls applied afterward). It also splits them by execution type into computational (deterministic and fast, running on CPU, such as tests, linters, and type checkers) and inferential (semantic analysis and LLM-as-judge, slower, costlier, and less deterministic). Mentioning this shows you followed the topic to its latest version. In the memo she also flagged a gap: OpenAI's write-up does not address verification of functionality and behavior, which is exactly what the evaluation section below fills.

OpenAI later turned this into a platform story. Codex as a platform argues that the open-source Codex harness powers the App, CLI, and IDE extension, and can be embedded in your own product through codex exec, the Codex SDK, and app-server. The timing can only be given as summer 2026, and the Codex CLI itself was open-sourced in April 2025, so this was not a "first release." The same post cites an ARC-AGI-3 example: retained reasoning and context compaction raised GPT-5.6 Sol's score by nearly three times, which is OpenAI's own account of its own model and harness.

What a harness contains, and recent research

Combining OpenAI's account, Böckeler's frameworks, and the site's own write-ups gives a checklist for interviews:

ComponentResponsibility
Context managementDynamic assembly, compression, memory, knowledge base
Tool orchestrationTool registry, selection, result handling
Sandbox and approval boundariesExecution isolation, least privilege, human confirmation for risky actions
Deterministic constraintsLinters, structural tests, schema validation
Feedback loopsReturning failure signals to the model so it can self-correct
ObservabilityTracing, logs, cost and latency
Session managementMulti-turn state, checkpoint and resume

The site's Advanced Harness Engineering Patterns covers the Tool Registry, Guard System, and Checkpoint-Resume. Anthropic's Harness Design and Phil Schmid on the Agent Harness offer two more angles, and The Model Is a Component, the Harness Is the System collects conclusions from several companies.

Preprints on the topic have appeared since 2026. Evidence strength varies, and an interview citation should say so:

PaperContentEvidence strength
Harness Engineering for Agentic AI Coding ToolsExploratory study of 2,853 GitHub repos; context files dominate and AGENTS.md is becoming an interoperable formatMulti-author empirical study, marked as published at AIware 2026
Natural-Language Agent HarnessesSpecifying harnesses in natural languageMethod proposal
Agentic Harness EngineeringObservability-driven automatic evolution of harnessesMulti-author method paper
From Model Scaling to System ScalingArgues that scaling the harness matters as much as scaling the modelSingle author, position paper, not experimental evidence
Adapting the Interface, Not the ModelAdapting the harness interface at runtime; improved 116 of 126 model and environment settingsPage marks it as work in progress

The other "harness": evaluation frameworks

In evaluation, "harness" has a second meaning that you should keep apart. EleutherAI's LM Evaluation Harness is a unified framework that tests different models on the same code and inputs. Maiorano's LLM Readiness Harness turns evaluation into a deployment decision workflow, combining benchmarks, OpenTelemetry, and CI quality gates, and is a single-author preprint. Both are evaluation harnesses, a different thing from an agent's execution harness.

How to answer in an interview

Start with the definition: a harness is the engineering wrapped around a non-deterministic model, where humans design the environment, specify intent, and build feedback loops. Then explain composition with Böckeler's three categories or guides and sensors, and say up front that the first is her interpretation and she later changed frameworks. Finish by noting that "harness" has an execution sense and an evaluation sense, so the interviewer sees you can tell them apart.

Evaluation and Monitoring: The Part of the Harness Most Often Skipped

What a complete system needs

CapabilityWhat it doesWhy it matters
Multi-dimensional metricsTrack task success, policy compliance, groundedness, retrieval hit rate, cost, and p95 latency togetherReadiness is not one score
Mixed scoringDeterministic checks (valid JSON, PII detection), statistical metrics, LLM-as-judgeEach method has different blind spots
CI quality gateBlock the PR below threshold and show the diff in the PRTurns evaluation from a report into a decision
Tracing and observabilityComponent-level traces to see which pipeline stage failedAvoids unnecessary changes
Online continuous evaluationScore sampled production traffic and alertOffline sets lag the real distribution
Feedback into the datasetTag bad cases and add them to the eval setEvery incident becomes permanent protection

Maiorano's results give a concrete example of "not a single metric": on FiQA in an SLA-first scenario, gpt-4.1-mini led on readiness and faithfulness while gpt-5.2 paid a substantial latency cost. The same paper's ticket-routing experiments show regression gates consistently rejecting unsafe prompt variants. This is a single-author result, good for illustrating design thinking and not for generalizing.

The most common follow-up on mixed scoring is whether LLM-as-judge can be trusted. It carries biases, so calibrate against human labels, fix the rubric, and add deterministic checks for critical items. The site's How to Rigorously Compare an Agent Before and After a Change covers golden-set sizing, judge bias, and statistical tests, and Self-Reflection + LLM-as-Judge covers letting a model evaluate itself. On tooling, see Promptfoo, Braintrust, and Arize Phoenix. For observability, there is the Langfuse guide and Agent Observability: From OTel Traces to Catching Hallucinations, Tool Misuse, and Infinite Loops.

The risks of having no evaluation

Without evaluation, risk does not arrive as one explosion. It accumulates where nobody can see it:

  • Silent regression. You change a prompt to fix one problem and quietly break three other cases. It is the signature risk of LLM development, because outputs are non-deterministic and traditional assertions miss it.
  • Hallucination reaching users. Nobody measures groundedness, so wrong answers go straight to users.
  • Drift. A provider updates the model or users' questions shift, and yesterday's passing cases fail today. Only continuous evaluation reveals it.
  • Security and privacy. Prompt injection and leakage of sensitive data. The site's Agent Security: Prompt Injection and Trust Boundaries covers layered defenses.
  • Cascading failure. When output feeds other automation, one error amplifies into an operational incident, so high-risk flows need human confirmation.
  • Compliance and accountability. Without records and tests you cannot show how the system behaves at its limits.

How to answer in an interview

Open with six phrases: multi-dimensional metrics, mixed scoring, gates, tracing, online evaluation, and feeding failures back. Give each one sentence. For risks, make silent regression the headline, since it best explains why LLMs need regression tests like any software. If asked about judge reliability, answer with calibration, a fixed rubric, and deterministic checks.

Team Adoption: CI/CD and Evaluation Gates for LLM Development

Traditional CI assumes you can assert on output. LLM output spans too wide a range, so you compare against reference answers or ask another model to judge. The whole flow looks like this:

flowchart LR
  A["PR: changes prompt, model, RAG config, or agent logic"] --> B["Fast checks: format, schema, lint"]
  B --> C["Eval gate: golden set, deterministic checks, LLM-as-judge, red team"]
  C -->|"Below threshold"| X["Block PR, show diff"]
  C -->|"Pass"| D["Staging, shadow, or canary"]
  D --> E["Production: online eval, tracing, alerts"]
  E -->|"Bad cases fed back"| F["Golden set"]
  F --> C

Stage by stage, each one should answer what it blocks, how, and at what cost:

  1. Triggers. Not only code: prompts, model versions, RAG settings, agent logic, and tool descriptions all trigger the pipeline. Put them all under version control, including the evaluation dataset, or results cannot be reproduced.
  2. Fast checks. Valid JSON, schema validation, lint. Milliseconds to seconds, run on every commit. Böckeler's keep quality left means exactly this: cheap checks before integration, expensive ones (broader review, mutation testing) after.
  3. Evaluation gate. Run the frozen golden set with deterministic checks first and LLM-as-judge for semantics. Run each case several times and look at the pass rate rather than a single right-or-wrong. Set explicit thresholds, block the PR below them, and post the diff against main. Frameworks such as DeepEval, which works in a pytest style, or the Promptfoo mentioned above, fit here.
  4. Pre-production. Beyond staging, use shadow deployment (mirror traffic without returning results to users), canary, and A/B, plus human spot checks. This stage measures latency and cost.
  5. Production monitoring. Online sampled scoring, drift detection, anomaly alerts, and user feedback collection.
  6. Feeding back. Tag bad cases and add them to the golden set so every incident becomes a permanent test. The site's AI-Native SDLC Playbook L9 applies this idea to CI for agent configuration.

Model selection belongs in the same flow. Weight readiness by scenario (cost-first, risk-first, latency-first) instead of chasing a single top score.

A first step you can take tonight: collect 20 to 50 real cases as a golden set and wire up the simplest CI check, so that "changing a prompt gets tested" becomes true, then add dimensions gradually. The site's L9 article recommends starting at the same scale.

How to answer in an interview

Walk through four stages: development, evaluation gate, deployment (shadow, canary, A/B), and production monitoring, then add the step that closes the loop, feeding bad cases back. If asked how to test non-determinism, say sample several runs and look at pass rates, with thresholds set as ranges. If asked how to start, say 20 to 50 real cases and one simple CI check, not a full platform on day one.

The Takeaway

The trade-off across the three layers fits on one line: the further out you go, the larger the investment and the more failures you remove at once. Prompts are cheap and fast but have a low ceiling. Context decides whether the model has a chance of being right. The harness decides whether the system notices and contains the error afterward. Evaluation is part of the harness, and it is the only mechanism that lets changes to the other two layers be proven effective.

Common follow-upOne-line answer
Will prompt engineering become obsolete?The techniques remain foundational, but the focus has moved up to context and harness
Do role prompts really help?They help tone and depth, not factual accuracy in any consistent way
Do we still need context engineering as windows grow?Yes: longer is costlier and slower, and accuracy still degrades with length
Can LLM-as-judge be trusted?It is biased, so calibrate, fix the rubric, and add deterministic checks

Questions that keep showing up in public question banks

These questions come from seven public question banks (compared in post 11 of this series), keeping the ones that repeat across banks plus a few newer question types that map onto sections of this article. "Independent sources" only counts overlap between banks. It says nothing about how often a question appears in real interviews, and the amitshekhar and pallavi banks cite no sources, so this article does not use their company tags. Only question titles and links to where they appear are listed here, with no answers reproduced.

QuestionIndependent sourcesQuestion-bank linksWhere it fits in this article
How do you evaluate an LLM / RAG system? What is the taxonomy of evaluation methods, and which metrics do you use?4om · aeg · AIML · amitWhat a complete system needs
Explain few-shot learning and chain-of-thought prompting. When should you use CoT?3aeg · amit · ksLayer 1, Prompt: Make One Interaction Clear
How do you evaluate an agent? Compare trajectory evals and final-outcome evals; why can SWE-bench pass rates be misleading?3om · aeg · amit · palWhat a complete system needs
How do you version and manage prompts in production, and roll back behavior?3om · aeg · amit · palTeam Adoption: CI/CD and Evaluation Gates for LLM Development
Why do people say "evals are the moat"? What is vibes-based evaluation vs. a formal eval framework?2om · aegEvaluation and Monitoring: The Part of the Harness Most Often Skipped
How do you structure prompts for consistent structured output (JSON, XML)?2amit · ksLayer 1, Prompt: Make One Interaction Clear
What is context engineering, and how is it different from prompt engineering?2om · amitDefinition and boundary
What is context rot, and how does context compaction work in long-running agents?2om · amitHow architecture raises context quality
What is the context window, and what happens when you exceed it? What is the "lost in the middle" problem?2aeg · amitHow poor design breaks things
What matters more for an agentic coding tool like Claude Code: the model or the harness?2om · amit · palDefinition and role
Build the evaluation harness for a new frontier model release. What does it need to do?2om · palThe other "harness": evaluation frameworks
What is LLM-as-a-judge evaluation, and what are its known biases and limitations?2om · amit · palWhat a complete system needs
What is LLM observability? Design the observability stack for a production LLM application.2om · amit · palWhat a complete system needs
How do you evaluate and monitor a model in production, not just offline, and detect drift?2aeg · amit · palWhat a complete system needs
How do you detect and measure hallucination rate in production?2aeg · palThe risks of having no evaluation
Walk me through error analysis for 500 flagged production failures; how do you diagnose a chatbot whose accuracy dropped from 95% to 80% before retraining?2om · aegThe risks of having no evaluation
How do you wire evals into CI so prompt or model changes can't silently regress quality? How does CI/CD for AI applications differ from traditional CI/CD?2om · amit · palTeam Adoption: CI/CD and Evaluation Gates for LLM Development
How do you build a golden dataset and a regression test suite for AI applications?2aeg · amitTeam Adoption: CI/CD and Evaluation Gates for LLM Development
How do you implement A/B testing for LLM systems, and test a new model before full deployment (A/B, canary, interleaved, shadow)?2aeg · amitTeam Adoption: CI/CD and Evaluation Gates for LLM Development
What belongs in a repository instruction file such as AGENTS.md or CLAUDE.md for a coding agent, and what should stay out?1omWhat a harness contains, and recent research
What operational and business metrics matter for AI systems beyond accuracy?1aegWhat a complete system needs
Your new model version scores higher on every benchmark, but users say it got worse. Why does this happen, and how do you find the problem?1amit · palWhat a complete system needs
What are your testing strategies for non-deterministic outputs?1aegTeam Adoption: CI/CD and Evaluation Gates for LLM Development
Your new prompt scores 78% vs the old prompt's 74% on a 100-example eval. Do you ship it?1omTeam Adoption: CI/CD and Evaluation Gates for LLM Development

How sources are counted: amit and pal appear to be maintained by the same organization and share 26 near-verbatim questions, so together they count as 1 source; the two ks banks have the same author and count as 1; om, aeg, and AIML count as 1 each, so the maximum is 5. This section lists only question titles and links; see the original repos for answers.

The line-number links for amit, pal and aeg point to the main branch as of 2026-10-03 and can shift after those repos change; if a link lands on a different question, search the original file for the question text.

References

Prompt layer

Context layer

Harness and evaluation

Question banks