Contract engineering · Agents

An agent engineer who makes tool-using workflows survive production

If your team needs agents that take real actions, with durable state, scoped permissions and a recovery path when a step fails, you can contract me as an AI agent engineer. I join under your manager, start by writing down every tool action and who may approve it, and build the workflow, its evaluation and its runbook in your codebase.

This is a good fit if…

  • You have an agent demo that impressed people, and now it needs to touch real systems: tickets, orders, records, payments or infrastructure.
  • Runs fail halfway and nobody can tell what was done, what was not, or whether it is safe to retry.
  • Security has asked what the agent can do, on whose authority, and how it is stopped. Nobody has a written answer.
  • You have an engineering team and a manager to direct the work, and want senior agent experience inside it.

Look elsewhere if…

  • You want a governed agent delivered as a fixed outcome with acceptance criteria. Commission governed AI agent implementation instead.
  • An agent already runs in production and keeps failing. Agent reliability and recovery engineering is scoped for that diagnosis.
  • You only need tools exposed safely to an assistant over MCP. MCP integration with controlled tool access is narrower and faster.

What you get

Contract brief covering tool actions, state, evaluation and production ownership

  • A written inventory of every tool action the agent can take, with its permission scope, approval rule and failure handling.
  • Workflow state stored durably, so an interrupted run can be inspected, resumed or rolled back.
  • Irreversible actions gated behind human approval or hard limits enforced in code, not in the prompt.
  • An evaluation suite covering task completion and wrong-action cases, run before each release.
  • A runbook and named production owner for when the agent misbehaves.

Responsibilities I can own

  • Inventory tool actions with your engineers and security lead: what each does, whether it can be undone, and who may approve it.
  • Design and implement workflow state so every step is recorded, idempotent where possible and resumable.
  • Enforce permissions at the tool boundary with scoped credentials, not by instructing the model.
  • Build approval steps for consequential actions and a kill switch your operators control.
  • Create evaluation scenarios for success, partial failure, wrong tool choice and adversarial input.
  • Instrument runs so a single trace shows the model's decisions, tool calls and outcomes.
  • Write the runbook and pair with the engineers who will own the agent in production.

Stack fit

  • Python
  • TypeScript
  • Rust
  • LLM provider APIs
  • Model Context Protocol (MCP)
  • LangGraph or custom orchestration
  • Workflow and queue engines
  • PostgreSQL
  • OAuth and scoped credentials
  • OpenTelemetry tracing
  • Evaluation harnesses

Onboarding I need from you

  • Repository and staging access, including sandbox or test accounts for every system the agent will act on.
  • A named manager who owns the backlog, and a security contact who can approve permission scopes.
  • The owners of downstream systems, available to agree what the agent may and may not do there.
  • Written rules for what data the agent may read, store and send to a model provider.

Reporting

I report to your engineering or AI lead, work your sprint cadence, and send a weekly written update covering shipped changes, evaluation results and any change to what the agent is allowed to do.

How it runs

  1. 01

    Brief and fit check

    You describe the workflow, the systems it touches and who owns them. I reply with questions, a plain view on fit, and a checked availability window.

  2. 02

    Write the contract brief for the agent

    In the first weeks we agree the tool-action inventory, the state model, the evaluation scenarios and who owns the agent in production. Building starts once that is written down.

  3. 03

    Build in sandboxes, then behind approvals

    Tools are exercised against test accounts first. Consequential actions go live behind human approval, and limits are relaxed only on evidence from evaluation and real runs.

  4. 04

    Handover to the production owner

    Runbook, evaluation suite, traces and the remaining backlog handed to the named owner, who has already run the incident drill with me.

What needs to be in place

  • A specific workflow with a business owner, not a general wish for agents.
  • Sandbox or test access to every system the agent will act on.
  • A security or platform contact with authority to approve credential scopes.
  • A contract route agreed up front: direct, via your agency or a partner, and your IR35 or equivalent determination.

Not included

  • Fully autonomous operation of irreversible or regulated actions without a human approval step.
  • Promised task-completion rates. Evaluation reports what the agent does reliably and where it still fails.
  • Out-of-hours on-call support unless agreed separately in writing.
  • Ownership of your security policy. I implement to it and flag gaps; your security lead decides.
  • Substitution by another engineer. Any specialist help is named and approved by you first.

Beyond the agent demo

Agent demos are easy to build and hard to trust. The demo books a meeting, updates a ticket or drafts a refund, and everyone in the room is impressed. Then someone asks what happens if the refund call times out after the ledger was updated, or whether the agent could close the wrong customer’s ticket, and the honest answer is that nobody knows.

That gap is what this contract is for. You do not need more prompts. You need an engineer who treats the agent as a distributed system that happens to have a model making some of the decisions, and who builds the state, permissions and recovery paths that any such system needs before it touches production.

What the work involves

The work starts with a written account of what the agent is allowed to do, and only then moves to building it.

  • Tool actions. Every action is listed: which system, what it changes, whether it can be undone, what credential it uses and who approves it. This is usually the first time the business owner, the security lead and the downstream system owners have seen the agent’s powers in one place.
  • State. Each step of a run is recorded durably before and after it executes. A run interrupted by a timeout, a deploy or a model error can be inspected and either resumed or compensated. Steps that may be retried are made idempotent.
  • Permissions. Enforcement happens at the tool boundary with scoped credentials. Instructions in the prompt are a convenience, never the control.
  • Approvals. Consequential or irreversible actions pause for a human decision, with the context the approver needs on one screen.
  • Evaluation. Scenarios cover not just “did it complete the task” but “did it pick the wrong tool”, “did it stop when it should” and “did it resist instructions hidden in the data it read”.
  • Observability. One trace per run, showing model decisions and tool calls together, so an incident can be reconstructed in minutes.

The signature deliverable

The anchor document is a contract brief for the agent covering tool actions, state, evaluation and production ownership. Your team keeps it and updates it whenever the agent’s powers change. Illustrative example of the tool-action section:

Tool actionSystemReversibleControlOn failure
Read order historyOrder servicen/a (read)Scoped read token, customer ID filterRetry, then report
Issue refund under limitPaymentsYes, within 24 hoursHard limit in tool codeRecord state, no automatic retry
Issue refund over limitPaymentsYes, within 24 hoursHuman approval in support consoleHold run until decision
Close ticketHelpdeskYesAllowed only after customer confirmation stepLeave open, flag to agent queue

Illustrative example showing the format, not a record from a client engagement.

How acceptance is judged

Your manager accepts changes through normal review. Each release runs the evaluation suite, including wrong-action scenarios, and the results go in the pull request. New tool actions are not enabled in production until the inventory is updated and the relevant system owner has agreed. Acceptance at the end is concrete: the inventory matches the code, interrupted runs can be resumed or compensated in a drill, and the named owner can explain and operate every control.

Ownership and handover

The agent, its permissions and its backlog belong to your team from day one. The production owner is named while the contract brief is being written, pairs on the build, and runs an incident drill with me before I leave. The written handover covers the inventory, state model, evaluation suite, runbook and remaining backlog.

When to choose something else

If you would rather commission a governed agent as a fixed outcome, use governed AI agent implementation. If an agent already runs and keeps failing, start with agent reliability and recovery engineering. If you only need tools exposed safely to an assistant, MCP integration is narrower. If the work is mostly model integration without actions, the contract LLM engineer role fits better.

Questions buyers ask

What separates an agent engineer from someone who builds agent demos?

Ask what happens when step four of seven fails. A demo builder will talk about prompts and frameworks. An engineer who has run agents in production will talk about where state is stored, which steps are idempotent, how a half-finished run is detected and resumed, and how permissions are enforced outside the model. Ask for evidence of those, not of framework familiarity.

Which agent framework do you use?

The one that fits your stack and team, including none. Orchestration frameworks help with some workflows and get in the way of others. What matters more is that state, permissions and approvals are explicit in your code, so your engineers can change frameworks later without rewriting the safety controls.

How do you stop an agent doing something it should not?

Permissions sit at the tool boundary, using credentials scoped to what the agent needs, so the model cannot exceed them whatever it is told. Irreversible actions go through human approval or hard limits. Every run is traced, and operators have a switch that stops the agent. The published Substrate Pattern describes this approach.

What contract basis do you work on?

Directly, through your preferred agency, or through a delivery partner, at the published day rate. Your organisation makes the IR35 or equivalent status determination. I can describe the working practices for your assessment, but I do not give tax advice.

Who owns the agent after you leave?

A named person in your team, agreed in the first weeks and written into the contract brief for the agent. They pair on the work, run the evaluation suite themselves and take part in at least one incident drill before handover. If nobody can be named, I will raise that early, because an agent without an owner should not go live.

Related engagements

Contract engineering · LLM applications

LLM engineer

We have a funded LLM backlog and need a senior engineer to implement it within our team. How would a personal contract be scoped?

You get:Named contractor remit, delivery backlog, access prerequisites and handover plan

Contract engineering · Retrieval

RAG engineer

We need embedded capacity to improve retrieval, permissions and answer evaluation in an existing team. What should the contract deliver?

You get:Retrieval backlog, relevance evaluation plan, access-control tasks and handover

AI delivery · Agent controls

Governed AI agents

Our agents can trigger real actions. How do we implement approvals and stop conditions instead of relying on prompts alone?

You get:Action-tier design, approval boundary, audit events and adversarial policy tests

AI delivery · Agent reliability

Agent reliability

Our agent workflows fail halfway through and retries may repeat external actions. How should recovery be designed?

You get:State model, idempotency strategy, replay tests and operator recovery runbook

Contract engineering · Enablement

AI enablement engineer

We need a hands-on engineer who can implement workflows and coach users rather than sell a training-only programme. How should the role be written?

You get:Embedded implementation-and-coaching remit, workflow backlog and adoption handover

Contract engineering · Customer deployment

Forward-deployed engineer

Our customer deployments need an engineer who can work across the product team and client environment. What responsibilities should the contract include?

You get:Customer deployment remit, integration dependencies and acceptance ownership

Further reading

Send an engineering brief

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics