Engineering enablement · Coding agents

Coding agents on real repositories, with scoped tasks, limited rights and clear stop rules

If your developers are starting to hand whole tasks to coding agents and you need rules before that spreads, you can commission a vendor-neutral rollout programme. I define which tasks agents may take, what they may access, which tests must pass and when an agent must stop, then pilot those rules on a real repository and measure the result.

This is a good fit if…

  • Developers already use agents that edit many files, run commands and open pull requests, in a terminal, an IDE or a hosted service.
  • You use more than one agent product, or expect to switch, and want team rules that do not depend on one vendor.
  • Agent-authored pull requests are arriving faster than reviewers can judge them.
  • Security wants to know what credentials, networks and environments agents can reach.

Look elsewhere if…

  • You use one coding assistant for completions and chat and want a team workflow. Use AI engineering enablement for software teams instead.
  • You are building agents into your own product for customers. Use production AI agents and agent infrastructure instead.
  • You want platform-specific setup for Claude Code or GitHub Copilot. Use those platform pages instead.

What you get

Bounded coding-agent workflow, permission model, test requirements and stop conditions

  • A task-scope policy: which kinds of work agents may take alone, which need a developer alongside, and which stay human-led.
  • A permission model for agent environments: repository, branch, credentials, network and commands.
  • Test requirements an agent task must meet before it starts and before its work is accepted.
  • Stop conditions that end an agent run and hand back to a person.
  • A measured pilot on one repository and a decision on wider adoption.

How it runs

  1. 01

    Inventory current agent use

    We find out which agents are in use, where they run, which credentials they hold and what they have already changed, without blame.

  2. 02

    Draft the rules with the team

    With senior engineers and security, we draft task scope, the permission model, test requirements and stop conditions for one repository.

  3. 03

    Pilot on real tasks

    Agents take real backlog items under the rules for at least two full sprint cycles. Every run outcome is logged; I review and adjust the rules weekly.

  4. 04

    Measure and decide

    We compare agent-authored work against the baseline for that repository, finalise the rules, and you decide whether to extend them.

What needs to be in place

  • A pilot repository with CI, branch protection and a test suite that covers the areas agents will touch.
  • An engineering owner for the rules and a security contact for the permission model.
  • Approved agent tools, or agreement on which ones the pilot may use.
  • Willingness to run agents in isolated environments rather than on unrestricted laptops, where the risk calls for it.

Not included

  • Agents merging their own pull requests or deploying to production.
  • Building a custom agent platform. This programme sets practices for agents you use.
  • Guaranteed productivity gains. The pilot reports measured outcomes, including tasks agents should not take.
  • Vendor selection or procurement.

From assistant to agent changes the risk

A coding assistant suggests; a developer decides line by line. A coding agent takes a task and works through it: reads the repository, edits many files, runs commands, retries when tests fail and opens a pull request. Several products now do this in terminals, IDEs and hosted environments, and developers adopt them quickly because they save real time on well-defined work.

What changes is where control sits. The developer is no longer approving each line. They are approving a task at the start and a pull request at the end, and everything between depends on what the agent was allowed to do. Without agreed rules, each developer improvises: one runs an agent with full shell access and their own cloud credentials, another lets it rewrite tests until they pass. Reviewers see the result, not the run.

This programme sets those rules once, for the team, in a way that does not depend on a single vendor.

The four rule sets

Task scope. Not all work suits an agent. We classify task types for your repository into three groups. Agent alone: well-tested, contained changes such as adding tests for existing behaviour, dependency updates with good coverage, or mechanical refactors. Agent with a developer: features behind flags, changes in moderately tested areas. Human-led: authentication, payments, data migrations, infrastructure and anything with weak test coverage.

Permission model. For each environment an agent runs in, we define: which repository and branches it can write to (its own working branches, never the main branch), which credentials it holds (none for production, scoped tokens for anything else), what network access it has, and which commands need approval. Where a product’s own settings cannot enforce a rule, the environment does: containers, scoped tokens, network allowlists, branch protection.

Test requirements. Before an agent task starts, the behaviour it changes must have tests, or writing them is a separate human-reviewed task first. Before its work is accepted, CI must pass on the changed code, and the agent must not have weakened or deleted tests that assert existing behaviour.

Stop conditions. Written before the run: stop after repeated CI failures, stop if edits leave the agreed paths, stop if the diff exceeds the size limit, stop if a credential or network call outside the model is needed. A stopped run hands back to the developer with its log.

These come from the same thinking as the Substrate Pattern I have published: constrain what an agent can do mechanically, and put human approval at the boundaries that matter. Through Neul Labs I build infrastructure for AI agents, so I approach this from the engineering side, not only the policy side.

The signature deliverable

You end with a bounded coding-agent workflow, a permission model, test requirements and stop conditions. Illustrative example of a stop-condition register:

ConditionDetected byAgent actionThen
CI fails three times on the same taskCI statusStop, attach logDeveloper decides
Edit outside agreed pathsDiff check in CIBlock mergeReviewer investigates
Test assertions removed or loosenedTest-diff checkBlock mergeReviewer investigates
Diff above size limitPull-request checkStopTask split
Needs a credential it lacksTool errorStopHuman decides, never auto-grant

Illustrative example. Conditions are set with your engineers and security.

How acceptance is judged

During the pilot every agent run is logged: task type, outcome (accepted, accepted after rework, rejected, stopped), review time and any defect found later. We compare agent-authored pull requests with the repository’s baseline for size, review rounds and escaped defects. The engineering owner accepts the rules. The pilot also shows which task types agents should not take, which is as useful as which they should.

Ownership and handover

Your engineering owner holds the rules and changes them by pull request. Security holds the permission model. The pull-request template and CI checks carry the rules into daily work, so they do not depend on memory.

Boundaries

For platform-specific setup, see Claude Code enablement or GitHub Copilot enablement. If the strain is in review and release rather than agent rights, see AI-assisted code review and release controls. If your engineers need to build and operate LLM features, see build internal AI engineering capability. The parent programme is AI engineering enablement for software teams.

Questions buyers ask

Why vendor-neutral rules rather than product settings?

Because teams use several agents and switch between them. The rules (what an agent may do, what it may touch, what must pass, when it stops) are the same whichever product runs the task. Each product's settings then implement those rules; where a product cannot enforce one, the environment or CI must.

What is a stop condition?

A rule that ends an agent run and hands back to a person, written before the run starts. Examples are repeated CI failures, edits outside the agreed paths, changes to tests that assert existing behaviour, a diff above the size limit, or a request for a credential it does not have.

Do agents need special tests?

They need tests that exist before the task. An agent judges completion by whether checks pass, so a task with no test for the behaviour is a task it can falsely complete. The test requirements say which tasks need tests written first, and that an agent may not weaken tests to make them pass.

Who is accountable for agent-authored code?

A named person: the developer who assigned the task, and the reviewer who approves it. The agent is a tool. That accountability rule is written into the workflow and the pull-request template.

Scope a workflow pilot

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics