The situation
Your agent has moved past answering questions. It can update a CRM record, send an email to a customer, issue a credit, close a ticket or change a configuration. In the demo, it behaves. The system prompt says it must confirm before refunds and must never email external addresses.
Then the security review asks the obvious questions. What stops it if a customer email contains instructions? What happens if it issues fifty credits in a minute? Who approved that change, and where is the record? Prompts cannot answer those questions, because the model treats them as guidance, and content the agent reads can contradict them.
This page is for commissioning the control layer: approvals and stop conditions implemented in the runtime, tested adversarially, and owned by your team.
What the work involves
Tier every action. I use a tiered model I have published for AI in financial services: read-only, advise (the agent drafts, a person approves), act under supervision (the agent acts, every action logged, attributed and reversible) and act autonomously (reserved for low-consequence, well-tested actions). Every action in the agent’s reach is placed in a tier with an approver, a reversal path and the evidence that must be recorded.
Enforce boundaries outside the model. Following the Substrate Pattern, the model proposes and the runtime decides. Each action has pre-conditions (is this user allowed, is this amount under the ceiling, is approval present) and post-conditions (did the effect match the request). If a check fails, the action does not happen. Default is deny.
Add stop conditions. Spend and volume ceilings, rate limits per action, anomaly triggers (an unusual burst of high-tier requests) and a kill switch that halts the agent cleanly and leaves state recoverable.
Make every side effect auditable. Structured events link each action to the requesting user, the agent session, the approver, the policy rule applied and the result.
Attack it. Adversarial tests try the realistic routes around policy: instructions injected into emails, documents and tool results; chains of low-tier actions that add up to a high-tier effect; replaying approvals; evading limits by splitting requests.
The signature deliverable
You receive the action-tier design, approval boundary, audit events and adversarial policy tests. Illustrative example of an action register extract:
| Action | Tier | Approver | Stop condition | Reversal |
|---|---|---|---|---|
| Look up order status | Read-only | None | Rate limit per user | Not needed |
| Draft customer reply | Advise | Support agent | None | Discard draft |
| Update delivery address | Supervised | None; audited | Only before dispatch | Restore previous address |
| Issue goodwill credit | Advise | Team lead | Daily ceiling per customer and in total | Reverse credit |
Illustrative example showing the format, not a client deliverable.
How acceptance is judged
Acceptance criteria are agreed after tiering: every action in scope is registered and enforced at its tier; the adversarial suite passes in CI, with no high-tier action executed without approval; stop conditions trigger correctly in tests; and audit events are complete for a sample of pilot actions. The security sponsor and product owner sign off.
Ownership and handover
The control layer, policy configuration and test suite live in your repositories. The platform owner and security sponsor receive the runbook: changing a tier, adding an action through the same review, operating the kill switch and reading the audit trail. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If tool access itself is uncontrolled, start with MCP integration with controlled tool access. If the agent fails halfway and repeats actions on retry, see agent reliability and recovery engineering. For the broader policy-to-control mapping in a regulated operation, see AI enablement for regulated operations. To measure the agent’s output quality alongside its controls, add production LLM evaluation.