AI delivery · Private inference

Run language models inside your boundary, with the trade-offs measured first

If data rules or operating constraints mean your AI workload cannot use a hosted model as it does today, you can commission me to assess and deliver a private or local deployment. I benchmark candidate open-weight models on your actual task, compare quality, hardware and operating burden against the hosted baseline, record the decision, and deploy the chosen option with access controls and an operations plan.

This is a good fit if…

  • Customer contracts, data residency rules or your security policy rule out sending this data to a public model API.
  • You need inference in a location without reliable connectivity: a factory, a ship, a branch, an air-gapped network or an edge device.
  • You are weighing self-hosting against a private cloud endpoint and want quality, hardware and staffing costs compared on your workload.
  • You have a platform owner who can operate the result.

Look elsewhere if…

  • The only goal is a lower bill, with no data or location constraint. Start with reduce LLM cost and latency; self-hosting is one option among several.
  • You want to adapt a small model to your task through training. Small-model fine-tuning and distillation covers that decision.
  • You need senior platform capacity inside your team for an ongoing serving roadmap. Hire an AI platform engineer instead.

What you get

Deployment decision record, workload benchmark, access controls and operations plan

  • A deployment decision record comparing hosted, private-cloud and self-hosted options against your constraints, with reasons.
  • A workload benchmark: quality on your task, latency, throughput and hardware utilisation for each candidate model and configuration.
  • A deployed inference service in your environment, sized for measured demand.
  • Access controls: who and what may call the model, with authentication, quotas and audit logging.
  • An operations plan covering monitoring, model updates, security patching, capacity and failure handling.

How it runs

  1. 01

    Brief and fit check

    You describe the workload, the constraint driving a private deployment and your infrastructure. I reply with questions and a view on fit before anything is signed.

  2. 02

    Constraints and baseline

    I write down the binding constraints with your security and data owners, and measure the current or hosted baseline on a representative task set.

  3. 03

    Benchmark candidates

    Candidate open-weight models, quantisation levels and serving configurations are run on your task set and target hardware, measuring quality, latency, throughput and resource use.

  4. 04

    Decide and deploy

    The decision record is reviewed with the sponsor; the chosen option is deployed with access controls, monitoring and capacity limits.

  5. 05

    Acceptance and handover

    Acceptance measures are checked in your environment and the platform owner takes over the service and operations plan.

What needs to be in place

  • A clear statement of the constraint (data, residency, connectivity, contract) from the owner who holds it.
  • A representative task set, or agreement to build a minimal one, handled under your data policy.
  • Access to target hardware or approval to provision it, and the infrastructure team that manages it.
  • A platform owner prepared to run a model service, including patching and upgrades.

Not included

  • Purchasing, installing or physically maintaining hardware.
  • Licence or legal review of model weights for your use; I flag licence terms, your counsel decides.
  • A promise that a private model will match a frontier hosted model's quality on every task.
  • Operating the service or providing on-call after handover unless agreed separately in writing.

The situation

Your AI workload works on a hosted model API, or would. Then a constraint arrives. A customer contract forbids sending their data to third-party model providers. A regulator or internal policy requires data to stay in a region or on your own infrastructure. The application must run on a factory floor, a vessel or an air-gapped network. Or a security sponsor simply will not approve the current data flow.

The answer is often assumed to be “run our own model”. That is sometimes right. It also changes almost everything: model quality, hardware, cost structure, staffing and who is woken up when it breaks. The question to answer first is how quality, hardware and operating burden compare for your workload.

What the work involves

Pin down the constraint. With your security and data owners, I write down exactly what is and is not allowed. “No public API” and “no data leaves our network” lead to very different designs: the first may be met by a private endpoint from a cloud provider, the second needs self-hosting.

Measure the baseline. On a representative task set from your workload, I measure what you get today (or would get from a hosted model): quality, latency and cost per task. This is the bar every private option is compared with.

Benchmark candidates on your task. Open-weight models of different sizes, quantisation levels and serving configurations are run on your target hardware. I measure task quality with the same scoring, latency percentiles under realistic concurrency, throughput, and memory and accelerator utilisation. Public leaderboard scores are a starting filter, not evidence.

Count the operating burden. Serving software, model updates, security patching, monitoring, capacity planning and on-call. A private deployment is a service your team must run.

Deploy with access controls. The chosen option is deployed in your environment behind authentication, per-client quotas and audit logging, so the model is not an open endpoint on the internal network.

The signature deliverable

You receive the deployment decision record, workload benchmark, access controls and operations plan. Illustrative example of a benchmark summary:

OptionTask qualityp95 latencyThroughput at target concurrencyOperating burden
Hosted model (baseline)92%2.1 sProvider-limitedLow
Private cloud endpoint, same model family92%2.4 sContractedLow to moderate
Self-hosted mid-size model, 8-bit87%1.6 sMeets target on two GPUsHigh
Self-hosted small model, 4-bit79%0.9 sMeets target on one GPUHigh

Illustrative example. Figures are placeholders showing the format, not results from a client.

The decision record also states what would change the answer, such as a provider offering the required data terms, a drop in volume or a smaller model closing the quality gap, so the choice can be revisited on evidence rather than habit.

How acceptance is judged

Acceptance measures are agreed after the benchmark: the deployed option meets the agreed quality floor and latency target on the task set in your environment; access controls reject unauthenticated and over-quota calls; monitoring and alerts work; and the platform owner has completed a model update and rollback using the operations plan. The sponsor and platform owner sign off.

Ownership and handover

The service, configuration and benchmark harness live in your infrastructure and repositories. The platform owner receives the operations plan and the decision record, so future upgrades are evaluated the same way. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If cost or latency is the real driver, start with reduce LLM cost and latency. If a smaller model needs adapting to reach quality, see small-model fine-tuning and distillation. For moving between hosted providers, see LLM provider and model migration. For the wider picture of inference and platform work, see AI infrastructure, inference and cost optimisation.

Questions buyers ask

Will an open-weight model be as good as the hosted model we use now?

On some tasks, close enough; on others, noticeably worse. That is exactly what the benchmark measures, on your task rather than public leaderboards. The decision record shows the quality gap alongside the constraint, so the sponsor can decide whether the trade is acceptable or whether a private cloud endpoint is the better route.

Is self-hosting cheaper?

Sometimes, at steady high volume. Often not, once hardware, idle capacity, engineering time, patching and on-call are counted. The decision record states the full operating burden, not just the hardware cost. If the constraint is data rather than cost, a private endpoint from a cloud provider may meet it with less operational work.

What hardware will we need?

It depends on the model size, quantisation, concurrency and latency target, which the benchmark establishes. I report measured utilisation and headroom for each option rather than a generic sizing table, so your infrastructure team can plan capacity and procurement.

Can it run on edge devices with no connection?

Yes, for suitably small models and bounded tasks. Edge deployment adds constraints on memory, power, update distribution and monitoring when devices are offline, and the operations plan covers how models are updated and how failures are noticed.

Who keeps the model up to date?

Your platform owner. The operations plan covers evaluating new model versions against the same task set before upgrading, applying security patches to the serving stack, and rolling back. Without that routine, a private deployment slowly falls behind on both quality and security.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics