AI delivery · Evaluation

A release gate that catches AI quality regressions before your customers do

If you need to change models, prompts or retrieval without breaking customer workflows, you can commission me to build a production evaluation system. I build a representative dataset from your real traffic, measure a baseline, locate where failures come from, and wire a release gate into your pipeline, then hand the harness and its maintenance routine to a named owner.

This is a good fit if…

  • You want to upgrade or switch a model and have no reliable way to know what will break.
  • Your team spends hours checking AI outputs by hand because nobody trusts them, and you cannot tell whether retrieval, prompts, tools or the task itself is the problem.
  • Releases of an LLM feature are decided by someone trying a few examples and saying it looks fine.
  • You have a QA lead or engineering manager who will own the evaluation set after launch.

Look elsewhere if…

  • You need a defensible research protocol or a publishable benchmark rather than a maintained production gate. That is research work, run through dipankar.cc.
  • You already know you are changing provider and need the migration itself managed. Use LLM provider and model migration, which builds on an evaluation like this.
  • The problem is clearly retrieval quality in an existing assistant. RAG quality rescue is more specific.

What you get

Engineering evaluation harness and remediation

  • A representative evaluation dataset drawn from real usage, with labelled expected behaviour and known hard cases.
  • A measured baseline for the current system, broken down by task type and failure source.
  • Automated and human-review scoring methods matched to each task, with their limits written down.
  • A release gate in your CI or deployment pipeline that blocks changes which regress agreed measures.
  • A remediation list of the highest-impact failures, with the fixes made where in scope.
  • A maintenance routine and a named owner, so the dataset keeps pace with the product.

How it runs

  1. 01

    Brief and fit check

    You describe the feature, the change you are worried about and how quality is judged today. I reply with questions and a view on fit before anything is signed.

  2. 02

    Failure analysis and dataset

    I sample real traffic, classify failures by source (retrieval, prompt, tool call, model, or an ill-defined task) and build the first labelled dataset with your reviewers.

  3. 03

    Harness and baseline

    Scoring is implemented per task type: exact checks where possible, rubric-based model grading where calibrated against human labels, human review where neither is reliable.

  4. 04

    Gate and remediation

    The harness is wired into your pipeline as a release gate, and the top failure causes are fixed or ticketed with evidence.

  5. 05

    Handover

    The owner runs a release through the gate with me, then takes over the dataset, thresholds and maintenance routine.

What needs to be in place

  • Access to a representative sample of real inputs and outputs, handled under your data policy.
  • Reviewers who know what a good answer looks like and can label for a few hours a week.
  • A deployment pipeline where a gate can run, or agreement on where it will run.
  • A named owner for the evaluation set after handover.

Not included

  • A promise that hallucinations or errors will reach zero. Evaluation measures and reduces them; it does not abolish them.
  • Publishable benchmarks or research claims about models in general.
  • Ongoing manual review of production outputs on your behalf.
  • Regulatory sign-off of the system; that stays with your accountable owners.

The situation

An LLM feature is in production. Customers use it. Now something needs to change: a newer model, a cheaper provider, a rewritten prompt, a different retrieval setup. Nobody can say with confidence what the change will break, because the only test is someone trying a dozen examples and deciding it looks fine.

The same gap shows up as review burden. Your team checks AI answers by hand because a wrong one costs too much, yet nobody can say whether the errors come from retrieval, the prompt, a tool call, the model, or a task that was never well defined. Without that diagnosis, every fix is a guess.

This page is for commissioning the missing piece: a production evaluation system with a release gate, built from your real usage and maintained by your team.

What the work involves

Start from real failures, not generic benchmarks. I sample real inputs and outputs and classify what goes wrong, by source. That classification usually changes the plan. If most failures are retrieval misses, prompt tuning will not help. If the task itself is ambiguous, no model will be reliable until the task is redefined.

Build a dataset that represents your users. Cases are drawn from production traffic, stratified by task type, with expected behaviour labelled by your reviewers. Hard and high-consequence cases are deliberately over-represented. Personal data is minimised or redacted under your policy.

Score each task the right way. Structured outputs get exact or schema checks. Classification gets agreement metrics. Open-ended answers get rubric-based grading only after the grader has been calibrated against human labels on your data. Where nothing automated is reliable enough, the harness routes a sample to people.

Make it a gate, not a report. The harness runs in your CI or deployment pipeline. A change that regresses agreed measures beyond tolerance does not ship without an explicit, recorded override. This is the same principle as mechanical guardrails in software delivery: constraints that run every time beat vigilance that runs when someone remembers.

The signature deliverable

You receive the engineering evaluation harness and remediation: dataset, scorers, baseline report, release gate and a remediation list with fixes applied where in scope. Illustrative example of a gate result:

Task typeCasesBaselineCandidateToleranceGate
Invoice field extraction18094% exact95% exact−1 ptPass
Policy question answering12081% rubric pass76% rubric pass−2 ptsBlock
Escalation detection900 missed1 missed0 missedBlock

Illustrative example. Figures are placeholders showing the format, not results from a client.

How acceptance is judged

We agree acceptance measures after the baseline: the dataset covers each task type in scope; scorers meet an agreed agreement rate with human labels; the gate runs in the pipeline and correctly blocks a seeded regression; and the remediation list is evidenced. Your QA lead or engineering manager signs off.

Ownership and handover

The dataset, scorers and gate live in your repositories. The owner receives the maintenance routine (adding production failures, re-labelling, threshold reviews) and runs one real release through the gate with me before handover. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If you are about to switch provider, LLM provider and model migration uses this kind of harness for shadow evaluation and rollback. If the pressure is cost or latency, reduce LLM cost and latency measures every optimisation against quality. If the problem is plainly retrieval in an existing assistant, start with RAG quality rescue. If you need a research-grade benchmark or protocol, that belongs at dipankar.cc.

Questions buyers ask

Can't we just use a model to grade the outputs?

Sometimes. Model-graded scoring is useful for open-ended answers, but only after it has been checked against human labels on your data. Where it disagrees with your reviewers too often, I use exact checks, structured comparisons or human review instead. The harness records which method scores which task and how well it agrees with people.

How large does the evaluation dataset need to be?

Large enough to cover each task type and the failures that matter, which is usually hundreds of cases rather than thousands at the start. Coverage of hard and high-consequence cases matters more than size. The maintenance routine adds new cases from production failures so the set stays representative.

Will this reduce the time our team spends checking AI answers?

It should show you where checking is necessary and where it is not, which is the basis for reducing it. Once failures are traced to their source and fixed, some review can be targeted rather than blanket. I do not promise a specific reduction before the baseline is measured.

Who keeps the dataset up to date after you leave?

A named owner on your side, usually in QA or the product engineering team. The handover includes a routine: when a production failure is reported, it is labelled and added to the set, and thresholds are reviewed at agreed intervals. Without that owner, an evaluation set goes stale within months.

How is this different from research benchmarking?

A production evaluation answers one question: is this change safe to release for our users? A research protocol answers whether a finding holds in general, with controls and statistical claims to match. If you need the second, that work is run through dipankar.cc.

Define a production quality gate

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics