Skip to content
All tags

#fault-tolerance

2 posts
aidebug

10% of Conversations Were Making Things Up: Debugging a Silent Retriever Failure

About 10% of conversations on our AI assistant platform randomly lost knowledge base tools — the agent hallucinated answers from training data instead. Root cause: a bare except Exception swallowed Elasticsearch connection failures during retriever initialization, silently skipping tool registration. Fix: retry + surface failures to system prompt + structured metadata tracking.

aideep-dive

Multi-Agent Error Propagation and Recovery: Borrowing Thirty Years of Weapons from Distributed Systems

At 99% accuracy per step over 100 steps, the error-free completion rate drops to just 36% -- error compounding is a structural problem, not something prompt tuning can fix. Distributed systems' supervisor trees, bulkheads, circuit breakers, sagas, and durable execution can be mapped almost one-to-one into agent orchestration. But LLMs introduce a failure class that traditional systems never had -- semantic errors that don't crash -- which require Inspector agents (recovering 96.4%) and redundancy voting (MAKER: one million steps with zero errors) to address.