What changes after launch
An AI workflow keeps changing after launch even if nobody touches its code. Model providers release new versions and retire old ones. Library updates change defaults. The documents a retrieval system draws on go stale. Usage grows, and with it the monthly bill. The kinds of input users send drift away from what the evaluation set covered. Each change is small; together they mean a workflow that was good at launch can be noticeably worse six months later, with nobody able to say when it happened.
Ordinary software support does not catch this, because the workflow keeps returning valid responses. What it needs is a regular, structured look at quality, dependencies and cost by someone who understands how these systems fail. This retainer provides that, within a stated capacity, without implying the round-the-clock support that a single practitioner cannot honestly offer.
What the work involves
Each review period follows the same cycle:
- Quality review. Re-run the evaluation suite, sample recent production inputs and outputs, and compare against the baseline. Drift is reported with examples, and new failure cases are added to the evaluation set.
- Dependency and provider check. Track model version changes, deprecation notices, provider term changes and library updates. Deprecations get a migration plan well before their date.
- Cost review. Model, retrieval and infrastructure cost against volume, with specific proposals such as caching, prompt reduction or routing simpler requests to smaller models, each measured for its quality effect before it ships.
- Incident follow-up. For incidents since the last review: root cause, fix, and a new test so it cannot recur silently.
- Improvement backlog. Remaining capacity goes on the backlog your owner prioritises.
The approach rests on mechanical checks, the principle behind Vibes Inside Guardrails: an evaluation suite and alerts catch regressions more reliably than anyone remembering to look.
The signature deliverable
You receive a bounded maintenance scope, review cadence, incident responsibilities and improvement backlog. Illustrative example of a maintenance scope summary:
| Element | Agreed terms |
|---|---|
| Workflows covered | Support-ticket triage agent; knowledge assistant for the operations team |
| Capacity | A stated number of days per month, not carried over beyond the next month |
| Review cadence | Quality and cost review every two weeks; dependency check weekly; written report monthly |
| Incident roles | Client operations owner is first responder and applies runbook fallbacks; I investigate the cause in working hours |
| Response | Acknowledgement within one working day; no out-of-hours cover |
| Excluded | New workflows, out-of-hours support, third-party outages, work beyond capacity without approval |
Illustrative example. Not taken from a client engagement.
How the arrangement is judged
The scope sets what is measured: evaluation scores against baseline, incidents and their recurrence, cost per unit of work, and backlog items completed. Each monthly report shows these alongside the capacity used. Your owner accepts each report, and the quarterly review decides whether to continue, adjust or end the retainer.
Ownership
Your organisation owns the workflow, its code, accounts and evaluation suite throughout. Everything I change goes through your review process, and every runbook is written so your team can act without me. A good maintenance arrangement should make your own team more capable over time, not more dependent.
When to choose something else
If the workflow is failing or unpredictable now, start with AI agent reliability and recovery engineering. If the main concern is cost, reducing LLM cost and latency is a focused engagement. If you want to set up a quality-measurement discipline across many AI features, see production LLM evaluation. If your team is taking over a system built by someone else, see AI system handover.