From assistant to agent changes the risk
A coding assistant suggests; a developer decides line by line. A coding agent takes a task and works through it: reads the repository, edits many files, runs commands, retries when tests fail and opens a pull request. Several products now do this in terminals, IDEs and hosted environments, and developers adopt them quickly because they save real time on well-defined work.
What changes is where control sits. The developer is no longer approving each line. They are approving a task at the start and a pull request at the end, and everything between depends on what the agent was allowed to do. Without agreed rules, each developer improvises: one runs an agent with full shell access and their own cloud credentials, another lets it rewrite tests until they pass. Reviewers see the result, not the run.
This programme sets those rules once, for the team, in a way that does not depend on a single vendor.
The four rule sets
Task scope. Not all work suits an agent. We classify task types for your repository into three groups. Agent alone: well-tested, contained changes such as adding tests for existing behaviour, dependency updates with good coverage, or mechanical refactors. Agent with a developer: features behind flags, changes in moderately tested areas. Human-led: authentication, payments, data migrations, infrastructure and anything with weak test coverage.
Permission model. For each environment an agent runs in, we define: which repository and branches it can write to (its own working branches, never the main branch), which credentials it holds (none for production, scoped tokens for anything else), what network access it has, and which commands need approval. Where a product’s own settings cannot enforce a rule, the environment does: containers, scoped tokens, network allowlists, branch protection.
Test requirements. Before an agent task starts, the behaviour it changes must have tests, or writing them is a separate human-reviewed task first. Before its work is accepted, CI must pass on the changed code, and the agent must not have weakened or deleted tests that assert existing behaviour.
Stop conditions. Written before the run: stop after repeated CI failures, stop if edits leave the agreed paths, stop if the diff exceeds the size limit, stop if a credential or network call outside the model is needed. A stopped run hands back to the developer with its log.
These come from the same thinking as the Substrate Pattern I have published: constrain what an agent can do mechanically, and put human approval at the boundaries that matter. Through Neul Labs I build infrastructure for AI agents, so I approach this from the engineering side, not only the policy side.
The signature deliverable
You end with a bounded coding-agent workflow, a permission model, test requirements and stop conditions. Illustrative example of a stop-condition register:
| Condition | Detected by | Agent action | Then |
|---|---|---|---|
| CI fails three times on the same task | CI status | Stop, attach log | Developer decides |
| Edit outside agreed paths | Diff check in CI | Block merge | Reviewer investigates |
| Test assertions removed or loosened | Test-diff check | Block merge | Reviewer investigates |
| Diff above size limit | Pull-request check | Stop | Task split |
| Needs a credential it lacks | Tool error | Stop | Human decides, never auto-grant |
Illustrative example. Conditions are set with your engineers and security.
How acceptance is judged
During the pilot every agent run is logged: task type, outcome (accepted, accepted after rework, rejected, stopped), review time and any defect found later. We compare agent-authored pull requests with the repository’s baseline for size, review rounds and escaped defects. The engineering owner accepts the rules. The pilot also shows which task types agents should not take, which is as useful as which they should.
Ownership and handover
Your engineering owner holds the rules and changes them by pull request. Security holds the permission model. The pull-request template and CI checks carry the rules into daily work, so they do not depend on memory.
Boundaries
For platform-specific setup, see Claude Code enablement or GitHub Copilot enablement. If the strain is in review and release rather than agent rights, see AI-assisted code review and release controls. If your engineers need to build and operate LLM features, see build internal AI engineering capability. The parent programme is AI engineering enablement for software teams.