AI delivery · Model migration

Change model or provider without silently changing what users get

If you need to change LLM provider or model, because of a deprecation, cost, data-location terms or capability, you can commission a gated migration. I build a compatibility matrix of every call, run the candidate in shadow against your current outputs, check tool calling and data handling, and cut over behind a flag with a tested rollback. You see parity evidence before users do.

This is a good fit if…

  • A model you depend on is being deprecated and you have a date by which every call must move.
  • You want to move to a cheaper, faster or self-hosted model and need evidence that quality holds before committing.
  • A customer contract or data-protection review requires processing in a different region or under different terms.
  • Your application uses tool calling, structured output or long prompts that behave differently across providers, and nobody has measured the differences.

Look elsewhere if…

  • The aim is to cut cost or latency across many levers, of which a model change is only one. Use reducing LLM cost and latency.
  • You are moving to models you host yourselves and need the hosting built. Use private and local LLM deployment.
  • You want to adapt a smaller model to your task by training it. Use small-model fine-tuning and distillation.

What you get

Compatibility matrix, shadow evaluation, rollback plan and data-handling review

  • A compatibility matrix covering every model call: features used, parameters, limits and the equivalent on the target.
  • A shadow evaluation comparing current and candidate outputs on real traffic or a representative replay.
  • Tool-calling and structured-output tests that pass on the target before cutover.
  • A data-handling review of the target: retention, training use, region and sub-processors, checked against your obligations.
  • A staged cutover behind a flag, with a rollback that has been rehearsed.

How it runs

  1. 01

    Inventory the calls

    I find every place your systems call a model, record the features each relies on, and build the compatibility matrix against the target provider's current documentation.

  2. 02

    Build the evaluation

    An evaluation set from real prompts and outputs, with task-specific checks: exact fields for structured output, correct tool selection, reference answers or rubric scoring where needed.

  3. 03

    Shadow and compare

    The candidate runs alongside the current model on replayed or mirrored traffic, without affecting users. Differences are reviewed and prompts adapted where needed.

  4. 04

    Cut over and stand by to roll back

    Traffic moves in stages behind a flag, with monitoring on the agreed measures and a rehearsed rollback. The old path is removed only after the agreed observation period.

What needs to be in place

  • Access to the code that calls models, and to request and response logs or a way to capture a representative sample.
  • Accounts or credentials for the target provider or model, with the contractual terms available.
  • An owner for each AI feature who can judge whether a changed output is acceptable.
  • Your data-protection requirements, or the person who holds them.

Not included

  • Promising identical outputs. Models differ; the migration measures differences and decides which are acceptable.
  • Legal advice on provider contracts or data-protection law. The data-handling review provides the technical facts.
  • Negotiating commercial terms with providers.
  • Ongoing monitoring after the observation period unless agreed separately.

Why model changes go wrong quietly

Changing an LLM provider or model looks like a one-line change: a new endpoint, a new model name, perhaps a new key. That is what makes it risky. Nothing breaks at deploy time. The extraction job still returns JSON, slightly more often missing a field. The agent still calls tools, occasionally the wrong one. The support assistant still answers, a little more confidently about things it should hedge. A contract clause about where data is processed is no longer true. None of this appears in error rates.

A migration needs gates: evidence, collected before users are exposed, that quality, tool behaviour and data handling have not changed in ways you would not accept. This engagement builds those gates and moves your traffic through them.

What the work involves

Compatibility matrix. I inventory every model call across your systems and record what each depends on: system prompts, tool or function calling, structured output modes, streaming, context length, temperature and other parameters, embeddings, rate limits and timeouts. For each, the matrix records the target’s equivalent according to the provider’s current documentation, and the gaps.

Evaluation set. Built from real prompts and outputs, sampled across features and difficult cases, with personal data handled under your policy. Checks are specific to each feature: field-level comparison for extraction, tool and argument correctness for agents, reference answers or rubric scoring for text.

Shadow evaluation. The candidate model runs alongside the current one on replayed or mirrored traffic. Outputs are compared automatically, disagreements are sampled for human review, and prompts are adapted where the target needs different phrasing. Embedding changes get special care, because a new embedding model usually means re-indexing and a retrieval evaluation, not just a swap.

Data-handling review. For the target provider and configuration: retention periods, whether inputs may be used for training, processing region, sub-processors and logging. I set these against your stated obligations and flag mismatches for your data-protection owner.

Cutover. Behind a feature flag, by feature or percentage of traffic, with monitoring on the agreed measures and a rollback that is rehearsed before it is needed.

My recent work includes AI agent infrastructure in regulated financial services and drop-in components for production AI stacks, where swapping a component without changing behaviour is the whole point.

The signature deliverable

You receive a compatibility matrix, shadow evaluation, rollback plan and data-handling review. Illustrative example of a compatibility matrix extract:

Call siteFeatures usedGap on targetShadow resultGate status
Invoice extractionStructured output, long contextSchema mode differs; needs adapterField agreement within agreed threshold after prompt changePassed
Support agentTool calling, five toolsParallel tool calls behave differentlyWrong tool chosen on a small share of casesBlocked: prompt and tool description revision
Search embeddingsEmbeddings for retrievalDifferent dimensionsRequires full re-index and retrieval evaluationScheduled separately
SummariesFree text, streamingNone materialRubric scores comparablePassed

Illustrative example. Not taken from a client engagement.

How acceptance is judged

Each feature owner agrees a parity threshold before the shadow run. A feature cuts over only when it passes its threshold, its tool or schema tests pass, and the data-handling review has no unresolved mismatch. Your platform lead accepts the migration when all features have moved, the rollback has been rehearsed, and the observation period has passed without breaching the agreed measures.

Ownership and handover

The evaluation set, the comparison harness, the adapters and the flag configuration live in your repositories. The harness is designed to be re-run for the next model change, which is rarely far away. Your named owner receives a runbook for adding a candidate model, running the shadow comparison and rolling back.

When to choose something else

If cost or latency is the real goal, reducing LLM cost and latency considers caching, routing and prompt changes alongside model choice. If you are moving to self-hosted models, see private and local LLM deployment. If you need an evaluation discipline across all your AI features, see production LLM evaluation.

Questions buyers ask

If the new provider has an OpenAI-compatible API, is this not just a configuration change?

The request format may be compatible while the behaviour is not. Tool-calling reliability, structured-output adherence, refusal behaviour, context limits, tokenisation and rate limits all differ. A compatible endpoint makes the switch easy to make; it does not make it safe. The shadow evaluation is what shows whether it is.

What counts as parity?

It is defined per feature before the comparison runs. For structured extraction it might be field-level agreement with the current outputs or with labelled answers; for tool use, the correct tool and arguments; for free text, a rubric scored by people or a checked model judge. Each feature owner signs off its own threshold.

Can we use a model to judge the outputs?

For some features, yes, if the judge is itself checked against a sample of human ratings first. For high-stakes outputs I recommend human review of the disagreements. The evaluation plan says which approach is used where and why, and the judge's agreement with human ratings is reported alongside its scores.

How is this priced?

As scoped delivery with acceptance criteria, shown in the engagement model on this page. Scope depends on the number of distinct model calls, whether tool calling or structured output is involved, and how much evaluation data already exists. The inventory step gives a firm basis for the estimate.

What if the candidate model is not good enough?

Then the evaluation shows where, and you decide: adapt prompts, keep certain features on the current model, choose another candidate, or stop. If the reason for moving is a deprecation date, we plan the fallback options early rather than discovering the gap at cutover.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics