AI delivery · Cost and latency

Lower LLM cost and latency, with any quality loss shown rather than hidden

If an AI feature is too slow or too expensive, you can commission me to reduce cost and latency without silently trading away quality. I profile the real workload, build a quality evaluation for the task, map the quality-cost-latency frontier of the realistic options, and implement the changes that hold up, with each one measured before it ships and the trade-offs left for you to decide.

This is a good fit if…

  • An LLM feature's inference bill is growing faster than its usage or revenue, and finance has started asking questions.
  • Users complain about slow responses, or latency blocks a use case such as live chat or inline suggestions.
  • Previous cost cuts (a cheaper model, shorter prompts) were made without measuring quality, and nobody is sure what was lost.
  • You have a platform or product owner, and ideally a finance sponsor, who will own the targets.

Look elsewhere if…

  • Data or location constraints force a private deployment. Use private and local LLM deployment.
  • You have already decided to switch provider and need the migration managed. Use LLM provider and model migration.
  • The bottleneck is unclear and may sit across infrastructure, data and serving. Start with AI infrastructure, inference and cost optimisation.

What you get

Workload profile, quality-cost-latency frontier and measured optimisation backlog

  • A workload profile: request types, volumes, token usage, retries, latency by stage and cost per completed task.
  • A quality evaluation for the task, so every option is judged on output quality as well as speed and cost.
  • A quality-cost-latency frontier showing which options are worth considering and which are dominated.
  • A measured optimisation backlog, with the accepted changes implemented and verified.
  • Dashboards that keep cost per task, latency and quality visible together after handover.

How it runs

  1. 01

    Brief and fit check

    You describe the feature, the bill or latency problem and the constraints. I reply with questions and a view on fit before anything is signed.

  2. 02

    Workload profile

    I instrument the feature and analyse real traffic: request mix, prompt and output sizes, retries, cache hit rates, sequential calls and where time and money go.

  3. 03

    Quality evaluation and frontier

    A task evaluation is built with your reviewers, and candidate changes are measured on quality, cost and latency together to produce the frontier.

  4. 04

    Implement and verify

    Agreed changes ship through your review process one at a time, each checked against the evaluation and in production monitoring.

  5. 05

    Handover

    Dashboards, the evaluation and the remaining backlog are handed to the owner, with the trade-off decisions recorded.

What needs to be in place

  • Access to request logs, usage and billing data for the feature.
  • A sample of real requests and outputs, handled under your data policy.
  • Reviewers who can judge output quality and help label an evaluation set.
  • An owner who can accept trade-offs between quality, cost and latency.

Not included

  • A promised saving or latency figure before the workload is profiled.
  • Negotiating pricing or commitments with model providers.
  • Shipping a cheaper configuration that fails the agreed quality tolerance without your explicit decision.
  • Cloud cost work unrelated to the AI feature.

The situation

The AI feature is a success, which is the problem. Usage has grown, and the inference bill has grown faster. Or users have started complaining that responses take too long, and a use case that needs fast responses is on hold. A CTO, platform lead or finance sponsor wants it fixed.

The obvious levers are easy to pull: switch to a cheaper model, cut the prompt, cap output length. Each will show a saving on the next invoice. What it does to answer quality is usually not measured, and the cost of that shows up later as support tickets, manual review or users quietly giving up.

This page is for doing it properly: reduce cost and latency, and keep quality visible at every step.

What the work involves

Profile the real workload. I instrument the feature and analyse real traffic rather than averages. Typical findings include a long system prompt resent on every call, retrieval returning far more context than the model uses, retries on validation failures that double the cost of some requests, sequential tool calls that could run in parallel, and identical requests that could be cached. Cost per completed task, including retries and failures, is the unit that matters.

Build the quality evaluation. With your reviewers, I build a task evaluation from real requests, with a scoring method checked against human judgement. If you already have one, I use it. Without it, cost work is guesswork.

Map the frontier. Candidate changes are measured on quality, cost per task and latency percentiles together: prompt and context reduction, prompt caching, response caching, batching, parallelism, structured outputs that cut retries, routing simple requests to smaller models, provider or model alternatives, and streaming for perceived latency. Options that lose on every dimension are discarded. The rest are real trade-offs.

Implement and verify, one change at a time. Each accepted change ships through your review process, is checked against the evaluation before release and is confirmed in production monitoring. Changing several things at once hides which one caused a regression.

The signature deliverable

You receive the workload profile, quality-cost-latency frontier and measured optimisation backlog. Illustrative example of a backlog extract:

#ChangeCost per taskp95 latencyQualityStatus
1Cache static system prompt and policy text−22%−0.3 sNo changeShipped
2Run account and order lookups in parallelNo change−1.4 sNo changeShipped
3Structured output schema to cut retries−9%−0.2 s+1 ptShipped
4Route simple queries to smaller model−31%−0.6 s−4 ptsOwner decision: rejected

Illustrative example. Figures are placeholders showing the format, not results from a client.

Row four is the point of the page: a large saving that costs quality is put in front of the owner as a decision, not shipped quietly.

How acceptance is judged

Targets are agreed after profiling: cost per task, latency at stated percentiles and a quality tolerance against the baseline. Acceptance means the implemented changes meet the targets on the evaluation and in production monitoring, with no change shipped outside the quality tolerance unless the owner decided it in writing.

Ownership and handover

All changes live in your repositories. The owner receives dashboards showing cost per task, latency and quality together, the evaluation for re-use on future changes, and the remaining backlog with its measurements. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If constraints require running models yourself, see private and local LLM deployment. If a small adapted model might replace a large one, see small-model fine-tuning and distillation. For switching providers with shadow evaluation and rollback, see LLM provider and model migration. If the bottleneck may be wider than one feature, start with AI infrastructure, inference and cost optimisation.

Questions buyers ask

Why not just switch to a cheaper model?

It might be the right answer, but it is the change most likely to lose quality without anyone noticing. Many features have cheaper wins first: trimming prompts, caching repeated work, removing unnecessary retries, running independent calls in parallel. Every option, including a cheaper model, is measured on the same quality evaluation.

What does a quality-cost-latency frontier mean in practice?

It is a chart and table of the options measured on all three dimensions. Options that are worse on every dimension than another are discarded. What remains are genuine trade-offs, such as slightly lower quality for much lower cost, which your owner decides. It prevents a single headline number hiding a loss elsewhere.

How do you know quality has not dropped in production?

The evaluation runs before each change ships, and a sample of production outputs is scored on the same rubric afterwards. Quality sits on the same dashboard as cost and latency, so a regression shows up next to the saving it came with.

Can you promise a percentage saving?

No. Savings depend on your workload, which is why profiling comes first. Some features have large, easy wins; others are already efficient and the honest finding is that the remaining options cost quality. You get the measured backlog either way.

Do the changes lock us into one provider?

Where possible they do not. Caching, prompt changes, routing and parallelism mostly sit in your own code. Where an option depends on a provider-specific feature, the backlog says so, so the lock-in is a decision rather than a side effect.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics