The capability gap
Calling a model API is easy. Most good software engineers have done it, and many teams have a working prototype. The gap shows up afterwards. Nobody knows how good the answers are, because there is no evaluation set. A provider changes a model version and quality shifts without anyone noticing. Costs per request are a guess. When a user reports a harmful or wrong answer, there is no trace to inspect and no runbook. The team can build the feature but cannot yet own it.
That ownership is what this programme builds. I have written about how engineering skills are shifting with AI: the work moves from writing every line towards specifying, evaluating and operating systems whose behaviour is probabilistic. Those skills are learned by doing them on something real.
What the work involves
The skills-to-work matrix. I assess the team against the skills production LLM work needs, and map each to a piece of real work in the chosen system. The skills typically include: prompt and configuration versioning, retrieval design, building and maintaining an evaluation set, tracing and observability, cost and latency budgeting, permission and safety boundaries, handling model or provider changes, and incident response for model behaviour.
The bounded internal system. We pick a system with real internal users and contained risk: an internal knowledge assistant, a document triage workflow, a support drafting tool. It must be small enough to ship in the programme and real enough to need evaluation, monitoring and incident handling.
The paired backlog. Each backlog item is led by a named engineer and tied to one or more skills. I pair with that engineer, review the work and run short working sessions when a topic first comes up. I do not take items myself unless the team agrees it is the fastest way to unblock them.
I have designed and built performance-critical parts of an AI agent stack in regulated financial services and build agent infrastructure through Neul Labs. The programme draws on that production experience, and on my published account of forward deployment engineering, about building AI systems that survive contact with production.
The signature deliverable
You end with a skills-to-work matrix, a paired implementation backlog and an ownership graduation checklist. Illustrative example of a matrix extract:
| Skill | Team today | Backlog item that exercises it | Lead engineer |
|---|---|---|---|
| Evaluation set and harness | Ad hoc manual testing | Build 150-case evaluation set from real queries; run in CI | Engineer A |
| Retrieval design | None | Chunking and metadata filters for policy documents | Engineer B |
| Tracing and cost budgets | Logs only | Request tracing with token cost per request | Engineer C |
| Model change management | None | Switch model version behind evaluation gate | Engineer A |
| Incident response | None | Runbook and simulated wrong-answer incident | Engineer D |
Illustrative example. The format, not a client’s team.
How acceptance is judged
The sponsor accepts the programme against the graduation checklist, not against attendance. The team must, without me leading: run and extend the evaluation; ship a prompt or model change through the evaluation gate; handle a simulated incident from detection to fix using the runbook; explain cost and latency against budget; and onboard a colleague to the system. The internal system’s own quality is measured by the evaluation the team built.
Ownership and handover
The system, the evaluation set, the runbook and the backlog belong to your team throughout. By the end, the engineering manager has a team that can take on the next LLM feature without external help, and a clear list of anything still to learn.
Boundaries
If developers need to use coding tools better, start with AI engineering enablement for software teams. If you need capacity rather than capability, hire a contract LLM engineer. If you need a production quality gate designed for an existing system, see production LLM evaluation. For rules on coding agents, see coding-agent rollout.