The situation
Your AI workload works on a hosted model API, or would. Then a constraint arrives. A customer contract forbids sending their data to third-party model providers. A regulator or internal policy requires data to stay in a region or on your own infrastructure. The application must run on a factory floor, a vessel or an air-gapped network. Or a security sponsor simply will not approve the current data flow.
The answer is often assumed to be “run our own model”. That is sometimes right. It also changes almost everything: model quality, hardware, cost structure, staffing and who is woken up when it breaks. The question to answer first is how quality, hardware and operating burden compare for your workload.
What the work involves
Pin down the constraint. With your security and data owners, I write down exactly what is and is not allowed. “No public API” and “no data leaves our network” lead to very different designs: the first may be met by a private endpoint from a cloud provider, the second needs self-hosting.
Measure the baseline. On a representative task set from your workload, I measure what you get today (or would get from a hosted model): quality, latency and cost per task. This is the bar every private option is compared with.
Benchmark candidates on your task. Open-weight models of different sizes, quantisation levels and serving configurations are run on your target hardware. I measure task quality with the same scoring, latency percentiles under realistic concurrency, throughput, and memory and accelerator utilisation. Public leaderboard scores are a starting filter, not evidence.
Count the operating burden. Serving software, model updates, security patching, monitoring, capacity planning and on-call. A private deployment is a service your team must run.
Deploy with access controls. The chosen option is deployed in your environment behind authentication, per-client quotas and audit logging, so the model is not an open endpoint on the internal network.
The signature deliverable
You receive the deployment decision record, workload benchmark, access controls and operations plan. Illustrative example of a benchmark summary:
| Option | Task quality | p95 latency | Throughput at target concurrency | Operating burden |
|---|---|---|---|---|
| Hosted model (baseline) | 92% | 2.1 s | Provider-limited | Low |
| Private cloud endpoint, same model family | 92% | 2.4 s | Contracted | Low to moderate |
| Self-hosted mid-size model, 8-bit | 87% | 1.6 s | Meets target on two GPUs | High |
| Self-hosted small model, 4-bit | 79% | 0.9 s | Meets target on one GPU | High |
Illustrative example. Figures are placeholders showing the format, not results from a client.
The decision record also states what would change the answer, such as a provider offering the required data terms, a drop in volume or a smaller model closing the quality gap, so the choice can be revisited on evidence rather than habit.
How acceptance is judged
Acceptance measures are agreed after the benchmark: the deployed option meets the agreed quality floor and latency target on the task set in your environment; access controls reject unauthenticated and over-quota calls; monitoring and alerts work; and the platform owner has completed a model update and rollback using the operations plan. The sponsor and platform owner sign off.
Ownership and handover
The service, configuration and benchmark harness live in your infrastructure and repositories. The platform owner receives the operations plan and the decision record, so future upgrades are evaluated the same way. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If cost or latency is the real driver, start with reduce LLM cost and latency. If a smaller model needs adapting to reach quality, see small-model fine-tuning and distillation. For moving between hosted providers, see LLM provider and model migration. For the wider picture of inference and platform work, see AI infrastructure, inference and cost optimisation.