AI delivery · Infrastructure and inference

Make your AI workload faster and cheaper, with quality measured at every step

If your AI application is too slow, too expensive or needs to run somewhere your current setup cannot, you can commission me to benchmark the real workload and improve it. I measure latency, throughput, quality and cost together, agree targets with you, implement the changes that hold up, and hand the benchmark and operating plan to your platform team.

This is a good fit if…

  • Your AI feature works but latency complaints or a growing inference bill are now a product or finance problem.
  • You are unsure whether the fix is caching, batching, a smaller model, a different provider, self-hosting or a rewrite of a hot path.
  • Throughput limits or rate limits are blocking growth, and you need a measured plan rather than a bigger instance.
  • You have a platform lead or FinOps owner who will own the result.

Look elsewhere if…

  • You need senior platform capacity inside your team on your own backlog. Hire an AI platform engineer on contract instead.
  • Quality, not speed or cost, is the main problem. Start with production LLM evaluation.
  • You need general cloud cost reduction unrelated to AI workloads. That is a different engagement, scoped separately.

What you get

Inference and platform optimisation engagement

  • A workload benchmark built from real traffic shapes, measuring latency percentiles, throughput, cost per task and quality together.
  • A clear view of where time and money go: model calls, retrieval, orchestration, network, or your own code.
  • Agreed targets, and implemented changes that meet them without an unmeasured quality drop.
  • A deployment and operating plan your platform team can run, including capacity and alerting.
  • A record of options considered and rejected, with the measurements behind each decision.

How it runs

  1. 01

    Brief and fit check

    You describe the workload, the symptom and the constraint (latency, bill, data location, rate limits). I reply with questions and a view on fit before anything is signed.

  2. 02

    Workload benchmark

    I instrument the request path, replay representative traffic and measure latency, throughput, cost and quality together, so no single number hides another.

  3. 03

    Options and targets

    I set out the credible options with expected gains and risks, and we agree targets and acceptance measures in writing.

  4. 04

    Implement and verify

    Changes ship through your review process, each verified against the benchmark and the quality evaluation before it is accepted.

  5. 05

    Handover

    Benchmark harness, dashboards, capacity notes and runbook transferred to your platform owner.

What needs to be in place

  • Access to the production request path or a faithful staging copy, and to cost and usage data.
  • A sample of real requests, handled under your data policy, that represents normal and peak traffic.
  • A quality measure for the task, or agreement to build a minimal one as part of the work.
  • A platform or engineering owner with authority to accept infrastructure changes.

Not included

  • A promised cost or latency figure before the workload is measured.
  • Purchasing or negotiating hardware, cloud commitments or provider contracts on your behalf.
  • Physical data-centre work or hardware installation.
  • Out-of-hours operation of the infrastructure after handover unless agreed separately in writing.

Which infrastructure route fits your constraint

Start here whenSignature outputMain acceptance measure
AI infrastructure and inference (this page) The bottleneck is unclear or spans several layersWorkload benchmark and implemented optimisationAgreed latency, throughput and cost targets met with quality held
Reduce LLM cost and latency One AI feature is too slow or expensive on hosted modelsQuality-cost-latency frontier and measured optimisation backlogEach change measured against a quality evaluation before it ships
Private and local LLM deployment Data or operating constraints require self-hosted or local inferenceDeployment decision record, workload benchmark, access controls, operations planQuality and latency on your workload, with the operating burden written down
Small-model fine-tuning and distillation You suspect a smaller, adapted model could do the taskBaseline comparison, rights-checked training plan and deployment decisionTask evaluation versus the prompting or RAG baseline, including a no-go
LLM provider and model migration You have decided to change provider or modelCompatibility matrix, shadow evaluation and rollback planShadow results within tolerance and a tested rollback

The situation

The AI feature works, and people use it. Then the symptoms arrive: responses take several seconds at the 95th percentile, the inference bill has grown faster than usage, provider rate limits throttle peak traffic, or a new customer’s data rules mean the current hosted setup is no longer allowed.

At that point the options multiply. Cache responses. Batch requests. Use a smaller model. Switch provider. Self-host. Fine-tune. Rewrite the orchestration layer. Each has advocates on your team, and each is easy to start and hard to judge, because the usual measurements look at one number at a time. A cheaper model that lowers the bill and quietly degrades answers is not a saving; it moves the cost to your users and your support team.

This page is the umbrella for that problem: benchmark the real workload, decide with evidence, implement what holds up.

What the work involves

Instrument the whole request path. I trace a request from arrival to response: input processing, retrieval, each model call, tool calls, post-processing and network. In agent and RAG systems, sequential calls and retries frequently cost more than the model itself.

Benchmark with real traffic shapes. Representative requests, at normal and peak concurrency, measured for latency percentiles, throughput, cost per completed task and task quality. Cost per task matters more than cost per token, because retries and failed attempts are part of what you pay.

Agree what quality means before touching anything. If the task has no quality evaluation, I build a minimal one with your reviewers first. Otherwise every saving is unverifiable, and the cheapest configuration always looks best.

Lay out the options honestly. Typical levers include prompt and context reduction, caching, request batching, parallelising independent calls, routing simple requests to smaller models, provider changes, self-hosted serving, quantisation, fine-tuned small models, and rewriting a measured hot path in a faster language. Each comes with an expected gain, a quality risk and an operating cost.

Implement and verify. Agreed changes ship through your review process. Each is accepted only when the benchmark shows the gain and the quality evaluation shows no regression beyond the agreed tolerance.

The signature deliverable

You receive the inference and platform optimisation engagement: the benchmark harness, the before-and-after measurements, the implemented changes and an operating plan. Illustrative example of a benchmark summary:

Configurationp95 latencyCost per task (relative)Task qualityDecision
Current6.8 s1.0091%Baseline
Parallel tool calls + prompt trim3.9 s0.8291%Adopt
Route simple requests to small model3.1 s0.5588%Reject: below tolerance

Illustrative example. Figures are placeholders showing the format, not results from a client.

How acceptance is judged

Targets are agreed after the benchmark, not before: latency percentile, throughput at a stated concurrency, cost per task and a quality floor. Acceptance means the targets are met on the benchmark and confirmed in production monitoring, and the platform owner signs off.

Ownership and handover

All changes live in your repositories and accounts. The platform owner receives the harness, dashboards, capacity notes and a runbook covering scaling, provider failures and rollback. I do the work personally; any specialist help is disclosed and approved by you first.

Choosing the narrower route

If one hosted-model feature is too slow or expensive, reduce LLM cost and latency is the focused version. If data or operating constraints require running models yourself, see private and local LLM deployment. If you want to test whether a smaller adapted model can do the task, see small-model fine-tuning and distillation. If the provider decision is already made, use LLM provider and model migration. For capacity under your own platform lead, hire an AI platform engineer.

Questions buyers ask

Why benchmark first instead of fixing the obvious problem?

Because the obvious problem is often not where the time or money goes. In AI applications, latency can sit in retrieval, sequential tool calls or serialisation rather than the model, and cost can come from retries or oversized prompts. A benchmark on real traffic shapes takes days and stops weeks of optimising the wrong layer.

Will you recommend rewriting things in Rust?

Only where the measurements justify it. Rust is valuable for specific hot paths such as tokenisation, parsing, vector operations and high-concurrency services, and I have built Rust components for AI stacks. Most AI latency problems are fixed with caching, batching, parallel calls or a smaller model, and those come first.

How do you make sure cost savings do not reduce quality?

Every change is run against a quality evaluation for the task as well as the cost and latency benchmark. A cheaper configuration that drops quality beyond the agreed tolerance is reported as a trade-off for you to decide, not shipped quietly.

Can this include moving to self-hosted models?

Yes, if the benchmark and your constraints support it. Self-hosting changes the cost structure and adds operational burden: GPUs, upgrades, monitoring and on-call. The private and local deployment route covers that decision in detail, including when it is the wrong move.

Who operates the system afterwards?

Your platform team. The handover includes the benchmark harness so they can re-run it after any change, dashboards for latency, cost and quality, capacity notes and a runbook. If you want bounded help after launch, ongoing maintenance is agreed separately.

Review an inference bottleneck

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics