AI delivery · Retrieval repair

Make your knowledge assistant find the right, current source

If your knowledge assistant misses documents, cites outdated versions or answers from the wrong source, you can commission a fixed-scope RAG rescue. I separate retrieval failures from generation failures on a labelled question set, add freshness and permission tests, then implement and evaluate the repairs. You get a measured before-and-after, the tests in your pipeline and a re-indexing runbook.

This is a good fit if…

  • Your internal knowledge assistant or support bot is live, and users report wrong, outdated or unsupported answers often enough that trust is falling.
  • Policies, prices or product documents change regularly and the assistant keeps quoting the superseded version.
  • You suspect the assistant can surface passages some users should not see, and need that checked and closed.
  • You want a defined result with acceptance criteria, rather than open-ended capacity under your own manager.

Look elsewhere if…

  • You want a retrieval engineer inside your team working your backlog. Hire a RAG and retrieval engineer on contract instead.
  • You have no retrieval system yet. Start with RAG and enterprise search engineering.
  • You need a generalisable finding about how models handle conflicting knowledge, rather than a repaired system. That is a research study, run through dipankar.cc.

What you get

Retrieval failure analysis, freshness tests, permission checks and evaluated repair

  • A labelled question set built with your experts, and a baseline showing how often answers fail and why.
  • Each failure classified as missing content, retrieval miss, ranking, staleness, permission leak or generation error.
  • Freshness tests that fail when the assistant cites a superseded or expired source.
  • Permission tests that prove restricted passages do not reach users without access.
  • Repairs implemented and measured against the same baseline, with remaining failures documented.

How it runs

  1. 01

    Collect the failures

    We gather real complaints and logged questions, and with two or three of your experts label the correct source and answer for a representative set.

  2. 02

    Diagnose

    I trace each failure through ingestion, chunking, indexing, retrieval, ranking and generation, and classify where it went wrong. The baseline report shows the split.

  3. 03

    Repair the largest causes

    I implement the fixes the diagnosis supports, such as metadata and version filters, re-indexing triggers, hybrid search, re-ranking or permission enforcement in the query, through your review process.

  4. 04

    Re-measure and hand over

    The same question set is re-run, freshness and permission tests are wired into CI, and the runbook for re-indexing and source changes goes to your owner.

What needs to be in place

  • Access to the RAG pipeline code, the index and a staging environment.
  • Question logs or a set of real user questions, with personal data handled under your policy.
  • Two or three subject experts who can label correct sources for a few hours.
  • The source systems' rules for versions, expiry and access, or someone who knows them.

Not included

  • A promised accuracy figure. The result is a measured change against your own baseline.
  • Rewriting or curating your source content. Content gaps are reported to their owners.
  • Replacing the whole RAG platform. Repairs are made to the system you have unless the diagnosis shows that is not viable.
  • Ongoing operation after handover unless agreed separately.

When the assistant stops being trusted

A retrieval-augmented assistant usually launches well. The demo questions work, the pilot users are pleased, and the system goes live across the organisation. Then the complaints start. It quoted last year’s travel policy. It said a product feature does not exist when the release notes say it does. It missed a document everyone knows is there. In the worst case, it showed a passage from an HR or legal folder to someone who should never have seen it.

Each complaint gets a quick fix, often a prompt change, and the quality drifts on. The underlying problem is that nobody can say, with evidence, where the answers fail. This engagement establishes that, repairs the largest causes, and leaves tests behind so the same failures cannot quietly return.

What the work involves

A labelled question set. I work with two or three of your subject experts to collect real questions, the passages that should answer them, and the passages that must not appear for particular users. A set of a few hundred well-chosen questions is usually enough to show the pattern.

Failure classification. For each failing question I look at what was retrieved and what the model did with it, and classify the cause:

  • Missing content. The answer is not in any indexed source. That is a content owner’s problem, and I report it to them.
  • Retrieval miss. The source is indexed but the right passage was not retrieved, often because of chunking, vocabulary or missing metadata.
  • Ranking. The passage was retrieved but ranked below the context limit.
  • Staleness. A superseded or expired version was retrieved instead of, or alongside, the current one.
  • Permission leak. A passage reached a user who lacks access to the source.
  • Generation. The right passage was in context and the answer was still wrong.

Freshness. Knowledge freshness has two parts: how quickly changes in the source reach the index, and whether the system knows which version is current. I check the re-indexing triggers and lag, add version and effective-date metadata where the sources provide it, and write tests that fail when a superseded source is cited.

Repair and re-measurement. I implement the fixes the diagnosis supports, in your codebase and through your review process, and re-run the same question set after each significant change, following the Vibes Inside Guardrails approach of mechanical checks rather than reviewer vigilance.

The signature deliverable

You receive retrieval failure analysis, freshness tests, permission checks and an evaluated repair. Illustrative example of a failure analysis summary:

Failure categoryShare of failing questions (baseline)Main cause foundRepairAfter repair
StalenessLargest shareOld policy PDFs never removed from index; no version metadataEffective-date metadata, superseded-version filter, deletion syncMost resolved; two sources lack dates, reported to owners
Retrieval missSecond largestTables in spec sheets split mid-rowStructure-aware chunking for tabular documentsSubstantially reduced
Permission leakSmall, but blockingGroup filter applied after retrievalAccess filter enforced in the index queryZero leaks on the permission test set
GenerationSmallModel ignores date caveats in contextPrompt and answer-format changePartly resolved; remainder documented

Illustrative example. Not taken from a client engagement.

How acceptance is judged

Acceptance criteria are set after the baseline, when we know the scale of each problem: for example, a target reduction in staleness failures on the question set, zero leaks on the permission tests, and a maximum re-indexing lag for named sources. Your search owner accepts the result against the same question set used for the baseline. Failures that remain are listed with their cause, so nothing is hidden behind an average.

Ownership and handover

The question set, the evaluation harness, the freshness and permission tests and the code changes all live in your repository. Tests run in CI on every retrieval change. Your named owner receives a runbook for re-indexing, adding a new source, changing access rules and investigating a reported bad answer.

When to choose something else

If you want a retrieval engineer working inside your team under your manager, see hire a RAG and retrieval engineer. If you are designing a retrieval system from the start, see RAG and enterprise search engineering. If the wider problem is measuring LLM output quality across several applications, see production LLM evaluation. If you need ongoing monitoring after the repair, see maintaining production AI workflows.

Questions buyers ask

How do you tell a retrieval problem from a model problem?

By looking at what was retrieved for each failed question. If the correct passage never reached the model, it is a retrieval, ranking, freshness or content problem. If it was in the context and the answer was still wrong, it is a generation problem. The labelled question set makes this a count rather than a guess.

How is this different from hiring your contract RAG engineer?

This is a commissioned outcome: a defined scope, acceptance criteria and me responsible for reaching them. The contract role buys my time under your manager and priorities. If your team wants to direct the retrieval work itself over months, the contract is the better fit.

Do we need to change vector database or model?

Usually not. Most failures I see come from ingestion, missing metadata, chunking, absent version handling or filters applied in the wrong place. Changing the embedding model or adding a re-ranker is considered only when the diagnosis shows ranking is the real cause, and the change is measured like any other.

How is this priced?

As scoped delivery with acceptance criteria, shown in the engagement model on this page. Scope depends on the number of sources, how access control works, and how many failure categories the diagnosis finds. A short scoping conversation about your sources and complaints comes first.

What about knowledge that conflicts between sources?

Within this engagement, conflicts are handled practically: version and authority metadata, source precedence rules and tests for the cases that matter to you. Researching how models detect or resolve conflicting knowledge in general is a separate validation study through the research practice.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics