Ongoing delivery · Maintenance

Keep a live AI workflow healthy without buying unlimited support

If an AI workflow or agent is live and nobody is watching for quality drift, provider changes or rising cost, you can commission a bounded maintenance retainer. For a stated number of days each month I run scheduled quality, dependency and cost reviews, act on agreed incidents in working hours, and work an improvement backlog your owner prioritises. Capacity is written down and not exceeded silently.

This is a good fit if…

  • An AI workflow, assistant or agent is in production and the team that built it has moved on to other work.
  • Model versions, provider terms and library dependencies keep changing, and nobody checks what each change does to output quality.
  • Monthly model and infrastructure cost has crept up and nobody is accountable for it.
  • Your internal owner runs the workflow day to day but wants a senior engineer on a regular cadence for review, fixes and improvements.

Look elsewhere if…

  • You need round-the-clock on-call cover or contractual response times outside working hours. Use your platform team or a managed service provider.
  • The workflow is not live yet, or is failing badly. Use production AI agents and agent infrastructure, or AI agent reliability and recovery engineering, first.
  • You need the system transferred to your own team. Use AI system handover and internal ownership.

What you get

Bounded maintenance scope, review cadence, incident responsibilities and improvement backlog

  • A written maintenance scope naming the workflows covered, the monthly capacity, working-hours response and what is excluded.
  • Scheduled reviews of output quality against the evaluation baseline, with drift reported before users notice it.
  • Dependency and provider changes tracked, with deprecations planned for rather than discovered.
  • Monthly cost reviewed against volume, with specific savings proposed and their quality effect measured.
  • An improvement backlog, prioritised by your owner, worked within the agreed capacity.

How it runs

  1. 01

    Onboarding review

    I review the workflow, its evaluation, monitoring, dependencies and costs, and write the maintenance scope with you. Gaps that must be fixed before maintenance makes sense are listed.

  2. 02

    Agree the cadence and roles

    We set the review schedule, who is first responder for each kind of incident, how I am contacted in working hours, and how unused or exceeded capacity is handled.

  3. 03

    Run the cycle

    Each period: quality review, dependency and provider check, cost review, incident follow-ups and backlog work, summarised in a short written report.

  4. 04

    Review the arrangement

    Each quarter we check whether the scope and capacity still fit. The retainer can be reduced, increased or ended on the agreed notice.

What needs to be in place

  • A live workflow with access to its code, deployment, monitoring and model provider accounts.
  • An evaluation set or baseline, or agreement to build one during onboarding.
  • A named internal owner who prioritises the backlog and is first contact for users.
  • Your incident process, so the retainer fits into it rather than beside it.

Not included

  • 24/7 or out-of-hours support, on-call duty or contractual response-time commitments.
  • Unlimited fixes or feature work. Work beyond the stated capacity is agreed and scoped separately.
  • Responsibility for outages of model providers, cloud platforms or other third parties.
  • Workflows not named in the maintenance scope.

What changes after launch

An AI workflow keeps changing after launch even if nobody touches its code. Model providers release new versions and retire old ones. Library updates change defaults. The documents a retrieval system draws on go stale. Usage grows, and with it the monthly bill. The kinds of input users send drift away from what the evaluation set covered. Each change is small; together they mean a workflow that was good at launch can be noticeably worse six months later, with nobody able to say when it happened.

Ordinary software support does not catch this, because the workflow keeps returning valid responses. What it needs is a regular, structured look at quality, dependencies and cost by someone who understands how these systems fail. This retainer provides that, within a stated capacity, without implying the round-the-clock support that a single practitioner cannot honestly offer.

What the work involves

Each review period follows the same cycle:

  • Quality review. Re-run the evaluation suite, sample recent production inputs and outputs, and compare against the baseline. Drift is reported with examples, and new failure cases are added to the evaluation set.
  • Dependency and provider check. Track model version changes, deprecation notices, provider term changes and library updates. Deprecations get a migration plan well before their date.
  • Cost review. Model, retrieval and infrastructure cost against volume, with specific proposals such as caching, prompt reduction or routing simpler requests to smaller models, each measured for its quality effect before it ships.
  • Incident follow-up. For incidents since the last review: root cause, fix, and a new test so it cannot recur silently.
  • Improvement backlog. Remaining capacity goes on the backlog your owner prioritises.

The approach rests on mechanical checks, the principle behind Vibes Inside Guardrails: an evaluation suite and alerts catch regressions more reliably than anyone remembering to look.

The signature deliverable

You receive a bounded maintenance scope, review cadence, incident responsibilities and improvement backlog. Illustrative example of a maintenance scope summary:

ElementAgreed terms
Workflows coveredSupport-ticket triage agent; knowledge assistant for the operations team
CapacityA stated number of days per month, not carried over beyond the next month
Review cadenceQuality and cost review every two weeks; dependency check weekly; written report monthly
Incident rolesClient operations owner is first responder and applies runbook fallbacks; I investigate the cause in working hours
ResponseAcknowledgement within one working day; no out-of-hours cover
ExcludedNew workflows, out-of-hours support, third-party outages, work beyond capacity without approval

Illustrative example. Not taken from a client engagement.

How the arrangement is judged

The scope sets what is measured: evaluation scores against baseline, incidents and their recurrence, cost per unit of work, and backlog items completed. Each monthly report shows these alongside the capacity used. Your owner accepts each report, and the quarterly review decides whether to continue, adjust or end the retainer.

Ownership

Your organisation owns the workflow, its code, accounts and evaluation suite throughout. Everything I change goes through your review process, and every runbook is written so your team can act without me. A good maintenance arrangement should make your own team more capable over time, not more dependent.

When to choose something else

If the workflow is failing or unpredictable now, start with AI agent reliability and recovery engineering. If the main concern is cost, reducing LLM cost and latency is a focused engagement. If you want to set up a quality-measurement discipline across many AI features, see production LLM evaluation. If your team is taking over a system built by someone else, see AI system handover.

Questions buyers ask

What happens if something breaks outside working hours?

Your own incident process handles it. The retainer does not include on-call cover, and the runbooks are written so your first responder can apply fallbacks, pause the workflow or roll back without me. I review the incident in the next working period and fix the cause within the agreed capacity.

What if a month needs more time than the retainer covers?

I tell you as soon as I can see it, with what the extra work is and why. You decide whether to approve additional days, defer lower-priority backlog items or accept the risk. Capacity is never exceeded silently, and you are never billed for time you did not approve.

Do you maintain systems you did not build?

Yes, after an onboarding review. Some systems need work before maintenance is sensible, for example adding an evaluation baseline or basic monitoring, and that is scoped separately so the retainer is not consumed by catch-up work. The onboarding review tells you which of these gaps exist before you commit to a monthly arrangement.

How is the retainer priced?

Monthly, for a stated capacity, shown in the engagement model on this page. The capacity is set from the onboarding review: the number of workflows, how often their dependencies change and the size of the improvement backlog you want worked.

How do we know the retainer is worth keeping?

Each period's report shows what was reviewed, what drifted, what was fixed and what changed in cost and quality. Each quarter we review whether the scope still fits. If the workflow is stable and your team can run the reviews themselves, reducing or ending the retainer is the right outcome.

Related engagements

Further reading

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics