When a RAG contract is the right buy
Most teams that ask for a “RAG engineer” are not starting from scratch. A knowledge assistant or internal search product already exists. It answers most questions, but users have stopped trusting it: it cites the wrong policy version, misses documents everyone knows exist, or (worse) surfaces a passage the user should never have been able to see. The team that built it is now busy with the next feature, and nobody owns retrieval quality as a discipline.
That is the situation this contract is for. You do not need a new system or a consultancy. You need a senior engineer who has done retrieval work before to sit inside your team, find out where answers actually fail, and fix it in your codebase, under your manager’s priorities.
What the work involves
Retrieval problems hide behind generation problems. When an answer is wrong, the cause is usually one of four things, and each needs a different fix:
- Nothing relevant was retrieved. Chunking split the answer across boundaries, the embedding model does not understand your domain vocabulary, or the document was never indexed.
- The right passage was retrieved but ranked too low to make it into the context window. That calls for re-ranking, hybrid lexical-plus-vector search, or metadata filters.
- The passage was retrieved but should not have been. Permission metadata was lost at ingestion, or filters are applied after retrieval rather than in the query.
- Retrieval was fine; generation was not. The model ignored or misread good context. This is a prompt or model problem, and retrieval changes will not help.
The first job is to tell these apart with evidence. I build a relevance set with two or three of your domain experts (real questions, the passages that should answer them, and passages that must never appear for a given user) and wire it into an evaluation harness your team can run on every change.
The signature deliverable
At the end of the contract you hold a retrieval backlog, a relevance evaluation plan, access-control tasks and a handover. Illustrative example of a backlog extract:
| # | Finding | Evidence | Proposed change | Owner |
|---|---|---|---|---|
| 1 | Policy answers cite superseded versions | 14 of 60 policy questions retrieve an older version first | Add effective-date metadata; filter to current version by default | Platform team |
| 2 | Contractor accounts can retrieve HR-only passages | 3 leaking passages found in permission test set | Enforce group ACLs in the vector query, not post-filter | Security + platform |
| 3 | Product codes split across chunks | Recall on part-number queries well below other categories | Structure-aware chunking for spec sheets | Me, paired with team |
Illustrative example. The numbers are placeholders showing the format, not results from a client.
How acceptance is judged
Your engineering manager sets priorities and accepts each change through your normal review. Every retrieval change is measured against the same relevance set before and after, and the result goes in the pull request. Permission findings get a separate test set that must pass with zero leaks before a change merges. At the end, the evaluation harness, the backlog and the handover are the measurable output: your team can show what improved and keep measuring after I leave.
Ownership and handover
The system, the backlog and the decisions stay with your team throughout. I pair on changes rather than working in a corner, so by the end at least one of your engineers can run the evaluation, interpret it and make the next retrieval change without me. The written handover covers the harness, the remaining backlog in priority order, known risks, and the runbook for re-indexing and access changes.
When to choose something else
If you would rather commission a fixed outcome with acceptance criteria and have me responsible for delivering it, use RAG quality rescue. If you are designing a new retrieval system, start with RAG and enterprise search engineering. If the requirement is wider LLM application work, not just retrieval, the contract LLM engineer role is the better brief. If you only need recurring senior input a few days a week, see part-time AI engineering capacity.