AI delivery · Agent controls

Agents that can act, inside approvals and stop conditions they cannot talk past

If your AI agents can change records, send messages or move money, you can commission me to implement the controls around them. I classify every action into tiers, enforce approval boundaries and stop conditions in code rather than prompts, emit audit events for every side effect, and prove the policy holds with adversarial tests, then hand the control layer to your platform and security owners.

This is a good fit if…

  • A tool-using agent is approaching production and can trigger real actions: updating records, emailing customers, issuing credits, changing configuration.
  • Current safety relies on the system prompt telling the model what not to do, and your security sponsor will not accept that.
  • You need least privilege, human approval for consequential actions and an audit trail that shows who authorised what.
  • You operate in a controlled or regulated setting where approvals and evidence must map to an agreed policy.

Look elsewhere if…

  • The integration layer itself is the gap: identity, scopes and logging for tool access. Start with MCP integration with controlled tool access.
  • Workflows fail halfway and retries repeat actions. That is a recovery problem; use agent reliability and recovery engineering.
  • You need an AI governance policy or board-level framework written, rather than controls implemented in a system. That is advisory work, scoped separately.
  • You want the agent to make regulated decisions with no human approver. I do not build that.

What you get

Action-tier design, approval boundary, audit events and adversarial policy tests

  • An action register: every action the agent can take, classified by tier, with its approver and reversal path.
  • Approval boundaries enforced by the runtime, so a tier that needs a human cannot proceed without one, whatever the model outputs.
  • Stop conditions such as spend ceilings, rate limits, anomaly triggers and a kill switch that halt the agent safely.
  • Structured audit events linking each side effect to the requesting user, the agent, the approver and the policy rule applied.
  • An adversarial policy test suite that runs in CI and shows the controls hold against injection and misuse.
  • A handover to named platform and security owners.

How it runs

  1. 01

    Brief and fit check

    You describe the agent, the actions it can take and the policy it must follow. I reply with questions and a view on fit before anything is signed.

  2. 02

    Action inventory and tiering

    With the product owner and security sponsor, I list every action, classify it (read-only, advise, act under supervision, act autonomously) and agree the approver and evidence for each tier.

  3. 03

    Implement the control layer

    Approval gates, scoped permissions, pre- and post-conditions, stop conditions and audit events are built into the runtime around the model, in your codebase.

  4. 04

    Adversarial testing

    Tests attempt to bypass the policy: injected instructions in inputs and tool results, approval skipping, chained low-tier actions with high-tier effects, limit evasion.

  5. 05

    Acceptance and handover

    Results are reviewed against the agreed criteria; platform and security owners take over the policy configuration, tests and runbook.

What needs to be in place

  • An agent in development or pilot, with access to its code and environment.
  • A security sponsor and product owner who can decide action tiers and approvers.
  • Any existing policy, risk classification or regulatory requirement the controls must reflect.
  • Approvers who are available to act on approval requests during the pilot.

Not included

  • Writing your organisation's AI policy or providing legal or regulatory advice.
  • Autonomous execution of regulated or high-stakes decisions without a human approver.
  • Penetration testing of your wider infrastructure beyond the agent's control layer.
  • Monitoring approval queues or operating the agent after handover unless agreed separately in writing.

The situation

Your agent has moved past answering questions. It can update a CRM record, send an email to a customer, issue a credit, close a ticket or change a configuration. In the demo, it behaves. The system prompt says it must confirm before refunds and must never email external addresses.

Then the security review asks the obvious questions. What stops it if a customer email contains instructions? What happens if it issues fifty credits in a minute? Who approved that change, and where is the record? Prompts cannot answer those questions, because the model treats them as guidance, and content the agent reads can contradict them.

This page is for commissioning the control layer: approvals and stop conditions implemented in the runtime, tested adversarially, and owned by your team.

What the work involves

Tier every action. I use a tiered model I have published for AI in financial services: read-only, advise (the agent drafts, a person approves), act under supervision (the agent acts, every action logged, attributed and reversible) and act autonomously (reserved for low-consequence, well-tested actions). Every action in the agent’s reach is placed in a tier with an approver, a reversal path and the evidence that must be recorded.

Enforce boundaries outside the model. Following the Substrate Pattern, the model proposes and the runtime decides. Each action has pre-conditions (is this user allowed, is this amount under the ceiling, is approval present) and post-conditions (did the effect match the request). If a check fails, the action does not happen. Default is deny.

Add stop conditions. Spend and volume ceilings, rate limits per action, anomaly triggers (an unusual burst of high-tier requests) and a kill switch that halts the agent cleanly and leaves state recoverable.

Make every side effect auditable. Structured events link each action to the requesting user, the agent session, the approver, the policy rule applied and the result.

Attack it. Adversarial tests try the realistic routes around policy: instructions injected into emails, documents and tool results; chains of low-tier actions that add up to a high-tier effect; replaying approvals; evading limits by splitting requests.

The signature deliverable

You receive the action-tier design, approval boundary, audit events and adversarial policy tests. Illustrative example of an action register extract:

ActionTierApproverStop conditionReversal
Look up order statusRead-onlyNoneRate limit per userNot needed
Draft customer replyAdviseSupport agentNoneDiscard draft
Update delivery addressSupervisedNone; auditedOnly before dispatchRestore previous address
Issue goodwill creditAdviseTeam leadDaily ceiling per customer and in totalReverse credit

Illustrative example showing the format, not a client deliverable.

How acceptance is judged

Acceptance criteria are agreed after tiering: every action in scope is registered and enforced at its tier; the adversarial suite passes in CI, with no high-tier action executed without approval; stop conditions trigger correctly in tests; and audit events are complete for a sample of pilot actions. The security sponsor and product owner sign off.

Ownership and handover

The control layer, policy configuration and test suite live in your repositories. The platform owner and security sponsor receive the runbook: changing a tier, adding an action through the same review, operating the kill switch and reading the audit trail. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If tool access itself is uncontrolled, start with MCP integration with controlled tool access. If the agent fails halfway and repeats actions on retry, see agent reliability and recovery engineering. For the broader policy-to-control mapping in a regulated operation, see AI enablement for regulated operations. To measure the agent’s output quality alongside its controls, add production LLM evaluation.

Questions buyers ask

Why not just instruct the model not to take dangerous actions?

Because instructions are inputs the model weighs, not rules it must obey, and injected content can override them. Prompts are still useful for behaviour, but control has to sit in code the model cannot change: permission checks, approval gates and limits that run whatever the model outputs.

Won't approvals make the agent too slow to be useful?

Only if everything needs one. Tiering is the point: read-only and easily reversible actions usually run without approval, while irreversible, financial or customer-facing actions wait for a person. We measure the approval load in the pilot and adjust tiers with evidence, not by removing controls when they become inconvenient.

Does this make our agent compliant with the EU AI Act or similar?

It implements technical controls and evidence that your compliance owners can map to their obligations, and the tiered model I use was designed with frameworks such as NIST AI RMF and the EU AI Act in mind. Whether a system is compliant is a determination for your accountable owners and advisers, not something I certify.

What happens when an approver rejects or ignores a request?

The action does not happen, and the agent receives a defined result it must handle: stop, ask for clarification or escalate. Unanswered requests expire after an agreed time. Both outcomes are audited, so you can see where approvals are bottlenecks.

Can you work with our existing agent framework?

Usually, yes. The control layer wraps tool and action execution, which every framework has, so it does not require a rewrite. Where a framework hides tool execution in a way that cannot be intercepted, I will say so in the assessment and propose the smallest change that makes control possible.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics