Why model changes go wrong quietly
Changing an LLM provider or model looks like a one-line change: a new endpoint, a new model name, perhaps a new key. That is what makes it risky. Nothing breaks at deploy time. The extraction job still returns JSON, slightly more often missing a field. The agent still calls tools, occasionally the wrong one. The support assistant still answers, a little more confidently about things it should hedge. A contract clause about where data is processed is no longer true. None of this appears in error rates.
A migration needs gates: evidence, collected before users are exposed, that quality, tool behaviour and data handling have not changed in ways you would not accept. This engagement builds those gates and moves your traffic through them.
What the work involves
Compatibility matrix. I inventory every model call across your systems and record what each depends on: system prompts, tool or function calling, structured output modes, streaming, context length, temperature and other parameters, embeddings, rate limits and timeouts. For each, the matrix records the target’s equivalent according to the provider’s current documentation, and the gaps.
Evaluation set. Built from real prompts and outputs, sampled across features and difficult cases, with personal data handled under your policy. Checks are specific to each feature: field-level comparison for extraction, tool and argument correctness for agents, reference answers or rubric scoring for text.
Shadow evaluation. The candidate model runs alongside the current one on replayed or mirrored traffic. Outputs are compared automatically, disagreements are sampled for human review, and prompts are adapted where the target needs different phrasing. Embedding changes get special care, because a new embedding model usually means re-indexing and a retrieval evaluation, not just a swap.
Data-handling review. For the target provider and configuration: retention periods, whether inputs may be used for training, processing region, sub-processors and logging. I set these against your stated obligations and flag mismatches for your data-protection owner.
Cutover. Behind a feature flag, by feature or percentage of traffic, with monitoring on the agreed measures and a rollback that is rehearsed before it is needed.
My recent work includes AI agent infrastructure in regulated financial services and drop-in components for production AI stacks, where swapping a component without changing behaviour is the whole point.
The signature deliverable
You receive a compatibility matrix, shadow evaluation, rollback plan and data-handling review. Illustrative example of a compatibility matrix extract:
| Call site | Features used | Gap on target | Shadow result | Gate status |
|---|---|---|---|---|
| Invoice extraction | Structured output, long context | Schema mode differs; needs adapter | Field agreement within agreed threshold after prompt change | Passed |
| Support agent | Tool calling, five tools | Parallel tool calls behave differently | Wrong tool chosen on a small share of cases | Blocked: prompt and tool description revision |
| Search embeddings | Embeddings for retrieval | Different dimensions | Requires full re-index and retrieval evaluation | Scheduled separately |
| Summaries | Free text, streaming | None material | Rubric scores comparable | Passed |
Illustrative example. Not taken from a client engagement.
How acceptance is judged
Each feature owner agrees a parity threshold before the shadow run. A feature cuts over only when it passes its threshold, its tool or schema tests pass, and the data-handling review has no unresolved mismatch. Your platform lead accepts the migration when all features have moved, the rollback has been rehearsed, and the observation period has passed without breaching the agreed measures.
Ownership and handover
The evaluation set, the comparison harness, the adapters and the flag configuration live in your repositories. The harness is designed to be re-run for the next model change, which is rarely far away. Your named owner receives a runbook for adding a candidate model, running the shadow comparison and rolling back.
When to choose something else
If cost or latency is the real goal, reducing LLM cost and latency considers caching, routing and prompt changes alongside model choice. If you are moving to self-hosted models, see private and local LLM deployment. If you need an evaluation discipline across all your AI features, see production LLM evaluation.