AI delivery · Model adaptation

Find out if a smaller, adapted model can do the job, and when it is not worth it

If you want to know whether a smaller fine-tuned or distilled model can meet your quality and latency needs, you can commission me to test it properly. I baseline prompting and retrieval first, check your rights to the training data, train and evaluate candidates on your task, and give a go or no-go decision that counts maintenance cost. On a go, I deploy and hand it over.

This is a good fit if…

  • A narrow, high-volume task (classification, extraction, routing, structured generation) runs on a large hosted model and costs or latency are a problem.
  • You need a model that runs privately or on modest hardware, and general small models fall short on your task.
  • Someone on the team proposes fine-tuning and you want an honest comparison with better prompting or retrieval before committing.
  • You have, or can produce, labelled examples of the task, and an ML or engineering owner for the result.

Look elsewhere if…

  • The task needs up-to-date or document-specific knowledge. That is usually a retrieval problem; see RAG and enterprise search engineering.
  • You only need lower cost or latency and have not tried caching, routing or prompt changes. Start with reduce LLM cost and latency.
  • You want to train a large foundation model from scratch. That is outside this work.

What you get

Baseline comparison, rights-checked training plan, task evaluation and deployment decision

  • A baseline comparison: large model with good prompting, retrieval where relevant, and off-the-shelf small models, all on the same task evaluation.
  • A rights-checked training plan: data sources, permission to use each for training, personal-data handling and the labelling approach.
  • Trained candidate models (fine-tuned or distilled from a larger teacher) evaluated on held-out data.
  • A go or no-go deployment decision that includes retraining frequency, drift monitoring and ownership cost.
  • Where it is a go, a deployed model with its evaluation and retraining pipeline handed to your team.

How it runs

  1. 01

    Brief and fit check

    You describe the task, volume, constraints and available data. I reply with questions and a view on fit before anything is signed.

  2. 02

    Task evaluation and baseline

    I build a task evaluation with your reviewers, then measure prompting, retrieval and off-the-shelf small models against it. If one of these is good enough, we stop here.

  3. 03

    Rights check and training plan

    Each data source is checked with its owner for permission to use it in training, and a plan covers data preparation, labelling, any teacher-generated data and held-out test sets.

  4. 04

    Train and evaluate

    Candidate small models are fine-tuned or distilled and evaluated on held-out data for quality, latency and cost per task against the baseline.

  5. 05

    Decision and handover

    The go or no-go is reviewed with the owner. On a go, the model is deployed with monitoring and a documented retraining routine.

What needs to be in place

  • A well-defined task with clear success criteria, and reviewers who can label examples.
  • Candidate training data, with an owner who can confirm it may be used for training.
  • Compute for training, in your cloud account or approved environment.
  • An ML or engineering owner who will maintain the model and its retraining pipeline.

Not included

  • Training on data you do not have the right to use, or on outputs whose provider terms prohibit it.
  • Legal advice on data rights or model licences; I flag questions, your counsel decides.
  • A promised quality or cost improvement before the baseline is measured.
  • Ongoing retraining after handover unless agreed separately in writing.

The situation

A large hosted model does a narrow task well: classifying support tickets, extracting fields, routing requests, producing structured summaries. It also does it slowly and expensively, at a volume that is starting to matter, or it cannot run where you need it to. Someone suggests fine-tuning a small model. Someone else says fine-tuning is a maintenance trap.

Both can be right. A small model adapted to one task can be fast, cheap and private. It can also be a model nobody knows how to retrain, built on data nobody had the right to use, that quietly degrades when inputs change. The question is whether a smaller model can meet your quality and latency requirements, and whether that is worth the added maintenance.

What the work involves

An evaluation before any training. I build a task evaluation with your reviewers: real inputs, expected outputs and a scoring method checked against human judgement. Everything that follows is measured on it.

An honest baseline. Before training anything, I measure the alternatives: the large model with careful prompting and examples, retrieval where the task depends on knowledge, and off-the-shelf small models. A good share of fine-tuning proposals end here, because a cheaper option already meets the bar. That is a useful result, not a failed project.

A rights-checked training plan. Each training source is checked with its owner: may it be used to train a model, does it contain personal data, and does any provider’s terms restrict using its outputs as training data? Data that fails the check is excluded, and the plan records why.

Train and compare. Candidate small models are fine-tuned on labelled examples or distilled from a larger teacher, with teacher outputs filtered and checked. Each is evaluated on held-out data for quality, latency and cost per task, beside the baseline.

Count the maintenance. Retraining frequency, drift monitoring, base-model upgrades and who owns them are part of the decision, not an afterthought.

The signature deliverable

You receive the baseline comparison, rights-checked training plan, task evaluation and deployment decision. Illustrative example of a baseline comparison:

ApproachTask qualityp95 latencyCost per 1,000 tasks (relative)Maintenance
Large hosted model, tuned prompt94%2.8 s1.00Low
Small off-the-shelf model, same prompt78%0.4 s0.06Low
Small model fine-tuned on 1,500 labelled examples92%0.4 s0.07Retrain quarterly; drift monitor

Illustrative example. Figures are placeholders showing the format, not results from a client.

How acceptance is judged

Acceptance thresholds are agreed after the baseline: the quality floor relative to the baseline, latency and cost targets, and the maintenance the owner is prepared to carry. A go requires the candidate to meet them on held-out data. A no-go is a complete, accepted outcome. On a go, deployment is accepted when the model meets the same thresholds in your environment and the owner has run the retraining pipeline once.

Ownership and handover

Models, training code, data lineage and the evaluation live in your accounts. The ML or engineering owner receives the retraining routine, drift monitoring and the criteria for deciding when to retrain or retire the model. I do the work personally; any specialist help is disclosed and approved by you first.

Boundaries

If the main question is where to run models, see private and local LLM deployment. If you want cheaper or faster inference without training anything, start with reduce LLM cost and latency. If you need a release gate for model changes generally, see production LLM evaluation. For broader ML pipelines and serving, see Python and machine learning engineering.

Questions buyers ask

Why test prompting and RAG first?

Because they are cheaper to build and much cheaper to maintain. A fine-tuned model has to be retrained when the task, data or base model changes, and someone must own that. If a well-prompted model or retrieval meets your bar, fine-tuning adds cost without benefit. The baseline makes that visible.

What is the difference between fine-tuning and distillation here?

Fine-tuning adapts a small model on labelled examples of your task. Distillation uses a larger teacher model to generate or label training data for the small model. Distillation can reduce labelling effort, but the teacher's output terms must allow it and its mistakes must be filtered out, so it still needs human-checked evaluation data.

How many labelled examples do we need?

It depends on the task's difficulty and variety. Narrow classification can work with hundreds of good examples; open-ended generation needs more. The baseline phase shows where the model struggles, so labelling effort goes to the cases that matter rather than to volume for its own sake.

What does a no-go look like?

A written decision showing the baseline, the best small-model result and the gap, with the reason it is not worth deploying: quality below the bar, savings too small to pay for maintenance, or data rights that do not allow training. You keep the evaluation set and data plan, which are useful whatever you do next.

Can the model run privately?

Yes. A small model adapted to one task is often the most practical way to run privately or on modest hardware. If private deployment is the driver, the private and local LLM deployment route covers serving, access controls and operations in more depth.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics