The situation
LLM features tend to arrive on a platform sideways. The first product team called a provider API directly with a key in an environment variable. The second did the same with a different SDK. A third wanted an open-weight model on a GPU node. Six months later there are a dozen paths to models, no shared rate limiting, no way to attribute token spend to a feature, and the first time a provider had an outage, three products failed in three different ways.
Your platform team knows how to run production infrastructure. What it is short of is senior time from someone who has built serving, observability and release controls specifically for model workloads, and who will do it inside your team’s practices rather than beside them.
What the work involves
Model workloads differ from ordinary services in a few ways that drive the backlog:
- Behaviour changes without a deploy. A provider updates a model, or someone edits a prompt in a config file, and output quality shifts. Release controls have to cover models and prompts, not just code.
- Cost is per request and variable. Token counts depend on inputs nobody controls. Cost attribution by feature and team is a platform concern, not a finance afterthought.
- Latency has a long tail. Streaming, retries and fallbacks need to be designed in, with timeouts set from measured distributions.
- Self-hosted serving is capacity planning. GPU memory, batching and cold starts behave differently from CPU services, and load testing is the only reliable guide.
The work usually starts with an inventory of every path to a model, then moves to a shared gateway or client, end-to-end instrumentation, and deployment controls, with self-hosted serving where it is in scope.
The signature deliverable
You end with a platform delivery backlog, deployment controls and an operational handover. The deployment controls are the part that prevents the next incident. Illustrative example:
| Control | What it prevents | Where enforced | Owner after handover |
|---|---|---|---|
| Model and prompt versions pinned per service | Silent behaviour change after provider updates | Gateway config in Git | Platform team |
| Evaluation gate on prompt or model change | Quality regression reaching users | CI pipeline | Owning product team |
| Canary rollout with automatic rollback on error rate | Full outage from a bad change | Deployment pipeline | Platform team |
| Per-team token budget with alert | Unattributed cost spikes | Gateway and cost dashboard | Platform team and finance partner |
Illustrative example showing the format, not a record from a client engagement.
How acceptance is judged
Your platform manager accepts each change through your normal change process. Platform components are measured against the baseline taken in the first weeks: latency distributions, error rates, cost attribution coverage. Deployment controls are accepted when a deliberate bad change in a non-production environment is caught and rolled back as designed. The operational handover is accepted when your on-call engineers have run the incident drill without me leading it.
Ownership and handover
Everything is defined in your infrastructure-as-code repositories, follows your conventions and is reviewed by your engineers. Each component has a named owner who paired on it. The handover covers architecture notes, runbooks for provider outage, latency spike, cost anomaly and quality regression, alert definitions, dashboards, and the remaining backlog in priority order.
When to choose something else
If one inference bottleneck or cost problem needs fixing as a scoped outcome, commission AI infrastructure, inference and cost optimisation or LLM cost and latency reduction. For a private model deployment delivered end to end, see private and local LLM deployment. If the application layer is the bottleneck rather than the platform, the contract LLM engineer role fits better.