AI delivery · Agent reliability

Agent workflows that survive timeouts without repeating real-world actions

If your agent workflows fail partway through, and retries risk sending, charging or updating something twice, you can commission me to design and implement recovery. I map the workflow's state and side effects, add idempotency and compensation for every external action, prove recovery with replay and fault-injection tests, and hand your operators a runbook for resuming or unwinding stuck runs.

This is a good fit if…

  • Agent runs sometimes stop halfway, after some external actions have happened and others have not.
  • You have seen, or fear, duplicate emails, double charges or repeated record updates after a timeout and retry.
  • Nobody can say exactly what state a failed run is in, so recovery means someone reading logs and fixing things by hand.
  • The agent is in or near production and you have an SRE or platform owner who will operate it.

Look elsewhere if…

  • The agent's answers are wrong, rather than its execution being unreliable. Use production LLM evaluation.
  • The gap is what the agent is allowed to do and who approves it. Use governed AI agent implementation.
  • You have no agent yet and need one built end to end. Start with production AI agents and agent infrastructure.

What you get

State model, idempotency strategy, replay tests and operator recovery runbook

  • An explicit state model for each workflow: steps, transitions, checkpoints and which steps have external side effects.
  • An idempotency key and deduplication strategy for every external action, so retries cannot repeat effects.
  • Compensation steps for actions that must be undone when a run is abandoned.
  • Durable checkpoints so a run can resume from the last safe point rather than starting again.
  • Replay and fault-injection tests in CI covering timeouts, crashes and partial failures at every step.
  • Observability and an operator runbook for finding, resuming, compensating or abandoning a stuck run.

How it runs

  1. 01

    Brief and fit check

    You describe the workflows, the failures seen and the systems they touch. I reply with questions and a view on fit before anything is signed.

  2. 02

    Failure and side-effect analysis

    I trace real failed runs, map each workflow's states and external effects, and classify every action as safe to retry, deduplicable or needing compensation.

  3. 03

    Implement recovery

    State persistence, idempotency keys, deduplication, compensation and resumption are built into your runtime or a durable execution layer your team can operate.

  4. 04

    Prove it with replay

    Fault-injection tests crash, time out and retry the workflow at every step and check that external systems end in a correct state.

  5. 05

    Acceptance and handover

    The SRE or platform owner reviews results against the agreed criteria and takes over the runbook, dashboards and tests.

What needs to be in place

  • Access to the agent's code, run logs and a non-production environment connected to test versions of the external systems.
  • Examples of failed or duplicated runs, if you have them.
  • Owners of the external systems who can confirm whether their APIs support idempotency keys or safe lookups.
  • A platform or SRE owner who will operate recovery after handover.

Not included

  • Changing the behaviour or APIs of third-party systems you do not control.
  • Improving the quality of the agent's reasoning or answers, beyond failures caused by execution.
  • Twenty-four-hour incident response for the agent unless agreed separately in writing.
  • Migrating your whole platform to a new workflow engine unless that is the agreed outcome of the analysis.

The situation

An agent workflow has six steps. It reads a request, looks up the customer, creates a ticket, issues a credit, emails the customer and closes the request. On Tuesday the email provider times out at step five. The framework retries the whole run. The customer gets two credits and, eventually, two emails. On Wednesday a deployment restarts the worker mid-run, and three requests are left half-done with nobody aware.

This is the most common way agent prototypes fail in production, and it has nothing to do with the model. It is a distributed-systems problem: external side effects, retries and partial failure. Adjusting prompts will not fix it.

This page is for commissioning the fix: an explicit state model, idempotent actions, tested recovery and a runbook your operators can follow.

What the work involves

Trace real failures. I start with failed and duplicated runs from your logs, if you have them, and reconstruct exactly what happened at each step. Often the cause is a retry policy set at the wrong level, such as retrying a whole run instead of a single call.

Make state explicit. Each workflow is modelled as states and transitions, with checkpoints persisted after every step that has an external effect. State is the primary unit of analysis: if you cannot say what state a run is in, you cannot recover it.

Classify every side effect. Each external action is one of: safe to repeat (a read), deduplicable (the target accepts an idempotency key, or we can check before acting) or compensable (it can be undone by a defined counter-action). Anything that is none of these is flagged for human confirmation on retry.

Implement recovery. Idempotency keys derived from the run and step, check-before-act lookups, compensation handlers, resumption from the last checkpoint, and bounded retries with backoff at the call level. Where many long-running workflows exist, I may recommend a durable execution engine; where a few exist, persisted state in your current stack is often enough.

Prove it. Fault-injection and replay tests kill the worker, time out each call and replay runs at every step, then check the external systems’ final state: no duplicates, no orphans, compensation applied where needed.

The signature deliverable

You receive the state model, idempotency strategy, replay tests and operator recovery runbook. Illustrative example of a side-effect classification:

StepExternal effectClassRecovery rule
Create ticketTicketing APIDeduplicableIdempotency key = run ID + step
Issue creditBilling APIDeduplicable, compensableCheck for credit with run reference before acting; reverse on abandon
Send emailEmail providerNot deduplicableRecord send before marking step done; on uncertain result, hold for operator
Close requestCRMSafe to repeatRetry with backoff

Illustrative example showing the format, not a client deliverable.

How acceptance is judged

Acceptance criteria are agreed after the analysis: every workflow in scope has a documented state model; every external action has a recovery rule; the replay and fault-injection suite passes in CI with no duplicate or orphaned effects; stuck runs appear in monitoring within an agreed time; and an operator, not me, recovers a seeded failure using only the runbook. The SRE or platform owner signs off.

Ownership and handover

All code, tests and dashboards live in your repositories and tools. The operator runbook covers finding stuck runs, resuming, compensating and abandoning them, and how to add a new step without breaking recovery. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If the agent’s actions need approval tiers and stop conditions, pair this with governed AI agent implementation. If tool access is the uncontrolled part, see MCP integration with controlled tool access. If you would rather have an engineer work this inside your team, hire an AI agent engineer. For the full agent build, see production AI agents and agent infrastructure.

Questions buyers ask

Can't we just stop retrying?

Then a transient timeout becomes a failed task and a person has to finish it by hand, often without knowing which steps completed. The fix is to make retries safe: each external action is deduplicated or checked before it repeats, and runs resume from a known checkpoint.

Do we need a durable execution platform?

Not necessarily. For a few workflows, persisted state, idempotency keys and a resumable runner in your existing stack are often enough. For many long-running workflows, a durable execution engine can be the right choice. The analysis recommends one with the operating cost stated; I do not introduce a platform nobody can run.

What if an external system cannot deduplicate requests?

Then the workflow checks before acting, for example by looking up whether the record or payment already exists using a key we control, and records the outcome before moving on. Where neither is possible, the step is marked for human confirmation on retry. The runbook states which steps are in that category.

How do operators know a run is stuck?

Each run's state is visible in a dashboard or query your team already uses, with alerts for runs that exceed expected duration or end in a failed state. The runbook explains, step by step, how to resume, compensate or abandon a run, and the audit trail shows what was done.

How is this different from improving the prompts?

Prompt changes can reduce how often the model makes mistakes, but they cannot make a timed-out API call safe to repeat. Duplicate actions are an execution problem: state, retries and side effects. This work is about the runtime, and it holds whatever model you use.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics