The situation
The AI feature works, and people use it. Then the symptoms arrive: responses take several seconds at the 95th percentile, the inference bill has grown faster than usage, provider rate limits throttle peak traffic, or a new customer’s data rules mean the current hosted setup is no longer allowed.
At that point the options multiply. Cache responses. Batch requests. Use a smaller model. Switch provider. Self-host. Fine-tune. Rewrite the orchestration layer. Each has advocates on your team, and each is easy to start and hard to judge, because the usual measurements look at one number at a time. A cheaper model that lowers the bill and quietly degrades answers is not a saving; it moves the cost to your users and your support team.
This page is the umbrella for that problem: benchmark the real workload, decide with evidence, implement what holds up.
What the work involves
Instrument the whole request path. I trace a request from arrival to response: input processing, retrieval, each model call, tool calls, post-processing and network. In agent and RAG systems, sequential calls and retries frequently cost more than the model itself.
Benchmark with real traffic shapes. Representative requests, at normal and peak concurrency, measured for latency percentiles, throughput, cost per completed task and task quality. Cost per task matters more than cost per token, because retries and failed attempts are part of what you pay.
Agree what quality means before touching anything. If the task has no quality evaluation, I build a minimal one with your reviewers first. Otherwise every saving is unverifiable, and the cheapest configuration always looks best.
Lay out the options honestly. Typical levers include prompt and context reduction, caching, request batching, parallelising independent calls, routing simple requests to smaller models, provider changes, self-hosted serving, quantisation, fine-tuned small models, and rewriting a measured hot path in a faster language. Each comes with an expected gain, a quality risk and an operating cost.
Implement and verify. Agreed changes ship through your review process. Each is accepted only when the benchmark shows the gain and the quality evaluation shows no regression beyond the agreed tolerance.
The signature deliverable
You receive the inference and platform optimisation engagement: the benchmark harness, the before-and-after measurements, the implemented changes and an operating plan. Illustrative example of a benchmark summary:
| Configuration | p95 latency | Cost per task (relative) | Task quality | Decision |
|---|---|---|---|---|
| Current | 6.8 s | 1.00 | 91% | Baseline |
| Parallel tool calls + prompt trim | 3.9 s | 0.82 | 91% | Adopt |
| Route simple requests to small model | 3.1 s | 0.55 | 88% | Reject: below tolerance |
Illustrative example. Figures are placeholders showing the format, not results from a client.
How acceptance is judged
Targets are agreed after the benchmark, not before: latency percentile, throughput at a stated concurrency, cost per task and a quality floor. Acceptance means the targets are met on the benchmark and confirmed in production monitoring, and the platform owner signs off.
Ownership and handover
All changes live in your repositories and accounts. The platform owner receives the harness, dashboards, capacity notes and a runbook covering scaling, provider failures and rollback. I do the work personally; any specialist help is disclosed and approved by you first.
Choosing the narrower route
If one hosted-model feature is too slow or expensive, reduce LLM cost and latency is the focused version. If data or operating constraints require running models yourself, see private and local LLM deployment. If you want to test whether a smaller adapted model can do the task, see small-model fine-tuning and distillation. If the provider decision is already made, use LLM provider and model migration. For capacity under your own platform lead, hire an AI platform engineer.