When the assistant stops being trusted
A retrieval-augmented assistant usually launches well. The demo questions work, the pilot users are pleased, and the system goes live across the organisation. Then the complaints start. It quoted last year’s travel policy. It said a product feature does not exist when the release notes say it does. It missed a document everyone knows is there. In the worst case, it showed a passage from an HR or legal folder to someone who should never have seen it.
Each complaint gets a quick fix, often a prompt change, and the quality drifts on. The underlying problem is that nobody can say, with evidence, where the answers fail. This engagement establishes that, repairs the largest causes, and leaves tests behind so the same failures cannot quietly return.
What the work involves
A labelled question set. I work with two or three of your subject experts to collect real questions, the passages that should answer them, and the passages that must not appear for particular users. A set of a few hundred well-chosen questions is usually enough to show the pattern.
Failure classification. For each failing question I look at what was retrieved and what the model did with it, and classify the cause:
- Missing content. The answer is not in any indexed source. That is a content owner’s problem, and I report it to them.
- Retrieval miss. The source is indexed but the right passage was not retrieved, often because of chunking, vocabulary or missing metadata.
- Ranking. The passage was retrieved but ranked below the context limit.
- Staleness. A superseded or expired version was retrieved instead of, or alongside, the current one.
- Permission leak. A passage reached a user who lacks access to the source.
- Generation. The right passage was in context and the answer was still wrong.
Freshness. Knowledge freshness has two parts: how quickly changes in the source reach the index, and whether the system knows which version is current. I check the re-indexing triggers and lag, add version and effective-date metadata where the sources provide it, and write tests that fail when a superseded source is cited.
Repair and re-measurement. I implement the fixes the diagnosis supports, in your codebase and through your review process, and re-run the same question set after each significant change, following the Vibes Inside Guardrails approach of mechanical checks rather than reviewer vigilance.
The signature deliverable
You receive retrieval failure analysis, freshness tests, permission checks and an evaluated repair. Illustrative example of a failure analysis summary:
| Failure category | Share of failing questions (baseline) | Main cause found | Repair | After repair |
|---|---|---|---|---|
| Staleness | Largest share | Old policy PDFs never removed from index; no version metadata | Effective-date metadata, superseded-version filter, deletion sync | Most resolved; two sources lack dates, reported to owners |
| Retrieval miss | Second largest | Tables in spec sheets split mid-row | Structure-aware chunking for tabular documents | Substantially reduced |
| Permission leak | Small, but blocking | Group filter applied after retrieval | Access filter enforced in the index query | Zero leaks on the permission test set |
| Generation | Small | Model ignores date caveats in context | Prompt and answer-format change | Partly resolved; remainder documented |
Illustrative example. Not taken from a client engagement.
How acceptance is judged
Acceptance criteria are set after the baseline, when we know the scale of each problem: for example, a target reduction in staleness failures on the question set, zero leaks on the permission tests, and a maximum re-indexing lag for named sources. Your search owner accepts the result against the same question set used for the baseline. Failures that remain are listed with their cause, so nothing is hidden behind an average.
Ownership and handover
The question set, the evaluation harness, the freshness and permission tests and the code changes all live in your repository. Tests run in CI on every retrieval change. Your named owner receives a runbook for re-indexing, adding a new source, changing access rules and investigating a reported bad answer.
When to choose something else
If you want a retrieval engineer working inside your team under your manager, see hire a RAG and retrieval engineer. If you are designing a retrieval system from the start, see RAG and enterprise search engineering. If the wider problem is measuring LLM output quality across several applications, see production LLM evaluation. If you need ongoing monitoring after the repair, see maintaining production AI workflows.