Beyond the agent demo
Agent demos are easy to build and hard to trust. The demo books a meeting, updates a ticket or drafts a refund, and everyone in the room is impressed. Then someone asks what happens if the refund call times out after the ledger was updated, or whether the agent could close the wrong customer’s ticket, and the honest answer is that nobody knows.
That gap is what this contract is for. You do not need more prompts. You need an engineer who treats the agent as a distributed system that happens to have a model making some of the decisions, and who builds the state, permissions and recovery paths that any such system needs before it touches production.
What the work involves
The work starts with a written account of what the agent is allowed to do, and only then moves to building it.
- Tool actions. Every action is listed: which system, what it changes, whether it can be undone, what credential it uses and who approves it. This is usually the first time the business owner, the security lead and the downstream system owners have seen the agent’s powers in one place.
- State. Each step of a run is recorded durably before and after it executes. A run interrupted by a timeout, a deploy or a model error can be inspected and either resumed or compensated. Steps that may be retried are made idempotent.
- Permissions. Enforcement happens at the tool boundary with scoped credentials. Instructions in the prompt are a convenience, never the control.
- Approvals. Consequential or irreversible actions pause for a human decision, with the context the approver needs on one screen.
- Evaluation. Scenarios cover not just “did it complete the task” but “did it pick the wrong tool”, “did it stop when it should” and “did it resist instructions hidden in the data it read”.
- Observability. One trace per run, showing model decisions and tool calls together, so an incident can be reconstructed in minutes.
The signature deliverable
The anchor document is a contract brief for the agent covering tool actions, state, evaluation and production ownership. Your team keeps it and updates it whenever the agent’s powers change. Illustrative example of the tool-action section:
| Tool action | System | Reversible | Control | On failure |
|---|---|---|---|---|
| Read order history | Order service | n/a (read) | Scoped read token, customer ID filter | Retry, then report |
| Issue refund under limit | Payments | Yes, within 24 hours | Hard limit in tool code | Record state, no automatic retry |
| Issue refund over limit | Payments | Yes, within 24 hours | Human approval in support console | Hold run until decision |
| Close ticket | Helpdesk | Yes | Allowed only after customer confirmation step | Leave open, flag to agent queue |
Illustrative example showing the format, not a record from a client engagement.
How acceptance is judged
Your manager accepts changes through normal review. Each release runs the evaluation suite, including wrong-action scenarios, and the results go in the pull request. New tool actions are not enabled in production until the inventory is updated and the relevant system owner has agreed. Acceptance at the end is concrete: the inventory matches the code, interrupted runs can be resumed or compensated in a drill, and the named owner can explain and operate every control.
Ownership and handover
The agent, its permissions and its backlog belong to your team from day one. The production owner is named while the contract brief is being written, pairs on the build, and runs an incident drill with me before I leave. The written handover covers the inventory, state model, evaluation suite, runbook and remaining backlog.
When to choose something else
If you would rather commission a governed agent as a fixed outcome, use governed AI agent implementation. If an agent already runs and keeps failing, start with agent reliability and recovery engineering. If you only need tools exposed safely to an assistant, MCP integration is narrower. If the work is mostly model integration without actions, the contract LLM engineer role fits better.