Contract engineering · Platform and inference

Senior platform capacity for model serving, observability and safe releases

If your platform team needs senior capacity for model serving, observability and deployment of LLM workloads, you can contract me as an AI platform engineer. I join under your platform engineering manager, work your infrastructure backlog in your repositories and pipelines, put deployment controls around models and prompts, and hand over runbooks your on-call engineers already use.

This is a good fit if…

  • Product teams are shipping LLM features faster than your platform can support them: each one calls providers differently, with its own keys and no shared monitoring.
  • You are moving some workloads to self-hosted or open-weight models and need serving, scaling and capacity planning done properly.
  • A model or prompt change has already caused an incident, and there is no way to roll it out gradually or roll it back.
  • Inference spend is rising and nobody can attribute it to a feature or team.

Look elsewhere if…

  • You want a specific inference bottleneck or cost problem diagnosed and fixed as a scoped outcome. Commission AI infrastructure, inference and cost optimisation instead.
  • You need a private model deployment delivered end to end. Private and local LLM deployment is scoped for that.
  • You need someone to run your production platform on call. This role builds and hands over; it is not an operations service.

What you get

Platform delivery backlog, deployment controls and operational handover

  • A shared path for model calls (gateway or library) with authentication, rate limits, timeouts and fallbacks in one place.
  • Traces, latency, error rates and token cost per feature and team, in your existing observability stack.
  • Model and prompt changes released like code: versioned, evaluated, rolled out gradually and reversible.
  • Serving for self-hosted models sized against measured load, with scaling and capacity alerts.
  • Runbooks and alerts your on-call engineers have rehearsed before handover.

Responsibilities I can own

  • Inventory current model usage: which services call which models, with what credentials, at what volume and cost.
  • Build or harden a model gateway or shared client with auth, quotas, timeouts, retries, fallbacks and cost attribution.
  • Set up serving for self-hosted models where in scope, with load testing, autoscaling and capacity alerts.
  • Instrument LLM calls end to end with traces, metrics and cost data in your observability stack.
  • Put deployment controls around models and prompts: versioning, evaluation gates, canary rollout and rollback.
  • Write runbooks for the failures that matter: provider outage, latency spike, cost anomaly, quality regression.
  • Pair with your platform engineers on each component so they own it before I leave.

Stack fit

  • Kubernetes
  • Terraform
  • Docker
  • AWS / Azure / GCP
  • vLLM and model servers
  • GPU scheduling
  • API gateways
  • OpenTelemetry
  • Prometheus / Grafana
  • Python
  • Rust
  • CI/CD pipelines

Onboarding I need from you

  • Access to infrastructure repositories, CI/CD and non-production environments through your normal joiner process.
  • A named platform engineering manager who owns the backlog and production change approval.
  • Read access to observability, cost and billing data for model usage.
  • Contacts in two or three product teams that consume model services.
  • Your change-management and incident processes, so my work follows them from the first change.

Reporting

I report to your platform engineering manager, work your sprint and change-approval cadence, and send a weekly written update covering shipped changes, platform metrics and operational risks.

How it runs

  1. 01

    Brief and fit check

    You describe the workloads, the platform and the gaps. I reply with questions, a plain view on fit, and a checked availability window.

  2. 02

    Inventory and baseline

    In the first weeks I map model usage, current latency, errors and cost by feature, and agree the platform backlog with your manager.

  3. 03

    Build the controls

    Gateway, observability, serving and deployment controls ship through your change process, highest-risk gaps first.

  4. 04

    Operational handover

    Runbooks, alerts and dashboards handed to your on-call engineers after a rehearsed incident drill, with the remaining backlog ranked.

What needs to be in place

  • An existing platform team and infrastructure-as-code practice I can work within.
  • A manager with authority over the platform backlog and production changes.
  • Budget approval for any new infrastructure, such as GPU capacity, before it is provisioned.
  • A contract route agreed up front: direct, via your agency or a partner, and your IR35 or equivalent determination.

Not included

  • Production on-call or 24/7 operations. I take part in incidents during agreed working hours only.
  • Promised cost reductions or latency figures. Changes are measured and reported as found.
  • Procurement of cloud or GPU capacity on your behalf.
  • Ownership of security policy. I implement to it and flag gaps; your security lead decides.
  • Substitution by another engineer. Any specialist help is named and approved by you first.

The situation

LLM features tend to arrive on a platform sideways. The first product team called a provider API directly with a key in an environment variable. The second did the same with a different SDK. A third wanted an open-weight model on a GPU node. Six months later there are a dozen paths to models, no shared rate limiting, no way to attribute token spend to a feature, and the first time a provider had an outage, three products failed in three different ways.

Your platform team knows how to run production infrastructure. What it is short of is senior time from someone who has built serving, observability and release controls specifically for model workloads, and who will do it inside your team’s practices rather than beside them.

What the work involves

Model workloads differ from ordinary services in a few ways that drive the backlog:

  • Behaviour changes without a deploy. A provider updates a model, or someone edits a prompt in a config file, and output quality shifts. Release controls have to cover models and prompts, not just code.
  • Cost is per request and variable. Token counts depend on inputs nobody controls. Cost attribution by feature and team is a platform concern, not a finance afterthought.
  • Latency has a long tail. Streaming, retries and fallbacks need to be designed in, with timeouts set from measured distributions.
  • Self-hosted serving is capacity planning. GPU memory, batching and cold starts behave differently from CPU services, and load testing is the only reliable guide.

The work usually starts with an inventory of every path to a model, then moves to a shared gateway or client, end-to-end instrumentation, and deployment controls, with self-hosted serving where it is in scope.

The signature deliverable

You end with a platform delivery backlog, deployment controls and an operational handover. The deployment controls are the part that prevents the next incident. Illustrative example:

ControlWhat it preventsWhere enforcedOwner after handover
Model and prompt versions pinned per serviceSilent behaviour change after provider updatesGateway config in GitPlatform team
Evaluation gate on prompt or model changeQuality regression reaching usersCI pipelineOwning product team
Canary rollout with automatic rollback on error rateFull outage from a bad changeDeployment pipelinePlatform team
Per-team token budget with alertUnattributed cost spikesGateway and cost dashboardPlatform team and finance partner

Illustrative example showing the format, not a record from a client engagement.

How acceptance is judged

Your platform manager accepts each change through your normal change process. Platform components are measured against the baseline taken in the first weeks: latency distributions, error rates, cost attribution coverage. Deployment controls are accepted when a deliberate bad change in a non-production environment is caught and rolled back as designed. The operational handover is accepted when your on-call engineers have run the incident drill without me leading it.

Ownership and handover

Everything is defined in your infrastructure-as-code repositories, follows your conventions and is reviewed by your engineers. Each component has a named owner who paired on it. The handover covers architecture notes, runbooks for provider outage, latency spike, cost anomaly and quality regression, alert definitions, dashboards, and the remaining backlog in priority order.

When to choose something else

If one inference bottleneck or cost problem needs fixing as a scoped outcome, commission AI infrastructure, inference and cost optimisation or LLM cost and latency reduction. For a private model deployment delivered end to end, see private and local LLM deployment. If the application layer is the bottleneck rather than the platform, the contract LLM engineer role fits better.

Questions buyers ask

What should the contractor own, and what should our team own?

I own building the components in the agreed backlog, to your standards and through your change process. Your team owns production, the change approval, on-call, budget and security policy from day one. Each component gets a named owner in your team before it ships, and that person pairs on it with me.

Is this LLMOps, MLOps or platform engineering?

For this role, it is platform engineering applied to model workloads: serving, gateways, observability, cost attribution and release controls for LLMs and other models. Classic MLOps topics such as training pipelines and feature stores are in scope where they sit on the same platform; if they are the main work, the Python and ML engineer role may fit better.

Do we need to self-host models?

Not necessarily. Many teams get most of the benefit from a shared gateway, observability and release controls in front of provider APIs. Self-hosting makes sense for data-residency, cost at steady high volume, or latency reasons. If it is in question, it goes on the decision list with a benchmark on your workload.

What contract basis do you work on?

Directly, through your preferred agency, or through a delivery partner, at the published day rate. Your organisation makes the IR35 or equivalent status determination. I can describe the working practices for your assessment, but I do not give tax advice.

Will you join our on-call rota?

During agreed working hours, I take part in incidents involving components I built, which is useful for learning how they fail. Out-of-hours on-call is not part of the default contract; if you need it, we agree it separately in writing. The aim is for your engineers to run the platform without me.

Related engagements

Contract engineering · LLM applications

LLM engineer

We have a funded LLM backlog and need a senior engineer to implement it within our team. How would a personal contract be scoped?

You get:Named contractor remit, delivery backlog, access prerequisites and handover plan

Contract engineering · Agents

AI agent engineer

We need someone to build tool-using workflows with reliable state, permissions and recovery, not another agent demo. What skills and scope fit?

You get:Contract brief covering tool actions, state, evaluation and production ownership

AI delivery · Infrastructure and inference

AI infrastructure

Our AI application is too slow and expensive. Who can benchmark the workload and improve cost without silently reducing quality?

You get:Inference and platform optimisation engagement

AI delivery · Private inference

Private LLM deployment

Our data or operating constraints require a private deployment. How do we compare quality, hardware and operational burden?

You get:Deployment decision record, workload benchmark, access controls and operations plan

Contract engineering · Enablement

AI enablement engineer

We need a hands-on engineer who can implement workflows and coach users rather than sell a training-only programme. How should the role be written?

You get:Embedded implementation-and-coaching remit, workflow backlog and adoption handover

Contract engineering · Customer deployment

Forward-deployed engineer

Our customer deployments need an engineer who can work across the product team and client environment. What responsibilities should the contract include?

You get:Customer deployment remit, integration dependencies and acceptance ownership

Further reading

Send an engineering brief

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics