The situation
The AI feature is a success, which is the problem. Usage has grown, and the inference bill has grown faster. Or users have started complaining that responses take too long, and a use case that needs fast responses is on hold. A CTO, platform lead or finance sponsor wants it fixed.
The obvious levers are easy to pull: switch to a cheaper model, cut the prompt, cap output length. Each will show a saving on the next invoice. What it does to answer quality is usually not measured, and the cost of that shows up later as support tickets, manual review or users quietly giving up.
This page is for doing it properly: reduce cost and latency, and keep quality visible at every step.
What the work involves
Profile the real workload. I instrument the feature and analyse real traffic rather than averages. Typical findings include a long system prompt resent on every call, retrieval returning far more context than the model uses, retries on validation failures that double the cost of some requests, sequential tool calls that could run in parallel, and identical requests that could be cached. Cost per completed task, including retries and failures, is the unit that matters.
Build the quality evaluation. With your reviewers, I build a task evaluation from real requests, with a scoring method checked against human judgement. If you already have one, I use it. Without it, cost work is guesswork.
Map the frontier. Candidate changes are measured on quality, cost per task and latency percentiles together: prompt and context reduction, prompt caching, response caching, batching, parallelism, structured outputs that cut retries, routing simple requests to smaller models, provider or model alternatives, and streaming for perceived latency. Options that lose on every dimension are discarded. The rest are real trade-offs.
Implement and verify, one change at a time. Each accepted change ships through your review process, is checked against the evaluation before release and is confirmed in production monitoring. Changing several things at once hides which one caused a regression.
The signature deliverable
You receive the workload profile, quality-cost-latency frontier and measured optimisation backlog. Illustrative example of a backlog extract:
| # | Change | Cost per task | p95 latency | Quality | Status |
|---|---|---|---|---|---|
| 1 | Cache static system prompt and policy text | −22% | −0.3 s | No change | Shipped |
| 2 | Run account and order lookups in parallel | No change | −1.4 s | No change | Shipped |
| 3 | Structured output schema to cut retries | −9% | −0.2 s | +1 pt | Shipped |
| 4 | Route simple queries to smaller model | −31% | −0.6 s | −4 pts | Owner decision: rejected |
Illustrative example. Figures are placeholders showing the format, not results from a client.
Row four is the point of the page: a large saving that costs quality is put in front of the owner as a decision, not shipped quietly.
How acceptance is judged
Targets are agreed after profiling: cost per task, latency at stated percentiles and a quality tolerance against the baseline. Acceptance means the implemented changes meet the targets on the evaluation and in production monitoring, with no change shipped outside the quality tolerance unless the owner decided it in writing.
Ownership and handover
All changes live in your repositories. The owner receives dashboards showing cost per task, latency and quality together, the evaluation for re-use on future changes, and the remaining backlog with its measurements. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If constraints require running models yourself, see private and local LLM deployment. If a small adapted model might replace a large one, see small-model fine-tuning and distillation. For switching providers with shadow evaluation and rollback, see LLM provider and model migration. If the bottleneck may be wider than one feature, start with AI infrastructure, inference and cost optimisation.