The situation
An agent workflow has six steps. It reads a request, looks up the customer, creates a ticket, issues a credit, emails the customer and closes the request. On Tuesday the email provider times out at step five. The framework retries the whole run. The customer gets two credits and, eventually, two emails. On Wednesday a deployment restarts the worker mid-run, and three requests are left half-done with nobody aware.
This is the most common way agent prototypes fail in production, and it has nothing to do with the model. It is a distributed-systems problem: external side effects, retries and partial failure. Adjusting prompts will not fix it.
This page is for commissioning the fix: an explicit state model, idempotent actions, tested recovery and a runbook your operators can follow.
What the work involves
Trace real failures. I start with failed and duplicated runs from your logs, if you have them, and reconstruct exactly what happened at each step. Often the cause is a retry policy set at the wrong level, such as retrying a whole run instead of a single call.
Make state explicit. Each workflow is modelled as states and transitions, with checkpoints persisted after every step that has an external effect. State is the primary unit of analysis: if you cannot say what state a run is in, you cannot recover it.
Classify every side effect. Each external action is one of: safe to repeat (a read), deduplicable (the target accepts an idempotency key, or we can check before acting) or compensable (it can be undone by a defined counter-action). Anything that is none of these is flagged for human confirmation on retry.
Implement recovery. Idempotency keys derived from the run and step, check-before-act lookups, compensation handlers, resumption from the last checkpoint, and bounded retries with backoff at the call level. Where many long-running workflows exist, I may recommend a durable execution engine; where a few exist, persisted state in your current stack is often enough.
Prove it. Fault-injection and replay tests kill the worker, time out each call and replay runs at every step, then check the external systems’ final state: no duplicates, no orphans, compensation applied where needed.
The signature deliverable
You receive the state model, idempotency strategy, replay tests and operator recovery runbook. Illustrative example of a side-effect classification:
| Step | External effect | Class | Recovery rule |
|---|---|---|---|
| Create ticket | Ticketing API | Deduplicable | Idempotency key = run ID + step |
| Issue credit | Billing API | Deduplicable, compensable | Check for credit with run reference before acting; reverse on abandon |
| Send email | Email provider | Not deduplicable | Record send before marking step done; on uncertain result, hold for operator |
| Close request | CRM | Safe to repeat | Retry with backoff |
Illustrative example showing the format, not a client deliverable.
How acceptance is judged
Acceptance criteria are agreed after the analysis: every workflow in scope has a documented state model; every external action has a recovery rule; the replay and fault-injection suite passes in CI with no duplicate or orphaned effects; stuck runs appear in monitoring within an agreed time; and an operator, not me, recovers a seeded failure using only the runbook. The SRE or platform owner signs off.
Ownership and handover
All code, tests and dashboards live in your repositories and tools. The operator runbook covers finding stuck runs, resuming, compensating and abandoning them, and how to add a new step without breaking recovery. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If the agent’s actions need approval tiers and stop conditions, pair this with governed AI agent implementation. If tool access is the uncontrolled part, see MCP integration with controlled tool access. If you would rather have an engineer work this inside your team, hire an AI agent engineer. For the full agent build, see production AI agents and agent infrastructure.