The situation
An LLM feature is in production. Customers use it. Now something needs to change: a newer model, a cheaper provider, a rewritten prompt, a different retrieval setup. Nobody can say with confidence what the change will break, because the only test is someone trying a dozen examples and deciding it looks fine.
The same gap shows up as review burden. Your team checks AI answers by hand because a wrong one costs too much, yet nobody can say whether the errors come from retrieval, the prompt, a tool call, the model, or a task that was never well defined. Without that diagnosis, every fix is a guess.
This page is for commissioning the missing piece: a production evaluation system with a release gate, built from your real usage and maintained by your team.
What the work involves
Start from real failures, not generic benchmarks. I sample real inputs and outputs and classify what goes wrong, by source. That classification usually changes the plan. If most failures are retrieval misses, prompt tuning will not help. If the task itself is ambiguous, no model will be reliable until the task is redefined.
Build a dataset that represents your users. Cases are drawn from production traffic, stratified by task type, with expected behaviour labelled by your reviewers. Hard and high-consequence cases are deliberately over-represented. Personal data is minimised or redacted under your policy.
Score each task the right way. Structured outputs get exact or schema checks. Classification gets agreement metrics. Open-ended answers get rubric-based grading only after the grader has been calibrated against human labels on your data. Where nothing automated is reliable enough, the harness routes a sample to people.
Make it a gate, not a report. The harness runs in your CI or deployment pipeline. A change that regresses agreed measures beyond tolerance does not ship without an explicit, recorded override. This is the same principle as mechanical guardrails in software delivery: constraints that run every time beat vigilance that runs when someone remembers.
The signature deliverable
You receive the engineering evaluation harness and remediation: dataset, scorers, baseline report, release gate and a remediation list with fixes applied where in scope. Illustrative example of a gate result:
| Task type | Cases | Baseline | Candidate | Tolerance | Gate |
|---|---|---|---|---|---|
| Invoice field extraction | 180 | 94% exact | 95% exact | −1 pt | Pass |
| Policy question answering | 120 | 81% rubric pass | 76% rubric pass | −2 pts | Block |
| Escalation detection | 90 | 0 missed | 1 missed | 0 missed | Block |
Illustrative example. Figures are placeholders showing the format, not results from a client.
How acceptance is judged
We agree acceptance measures after the baseline: the dataset covers each task type in scope; scorers meet an agreed agreement rate with human labels; the gate runs in the pipeline and correctly blocks a seeded regression; and the remediation list is evidenced. Your QA lead or engineering manager signs off.
Ownership and handover
The dataset, scorers and gate live in your repositories. The owner receives the maintenance routine (adding production failures, re-labelling, threshold reviews) and runs one real release through the gate with me before handover. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If you are about to switch provider, LLM provider and model migration uses this kind of harness for shadow evaluation and rollback. If the pressure is cost or latency, reduce LLM cost and latency measures every optimisation against quality. If the problem is plainly retrieval in an existing assistant, start with RAG quality rescue. If you need a research-grade benchmark or protocol, that belongs at dipankar.cc.