The situation
A large hosted model does a narrow task well: classifying support tickets, extracting fields, routing requests, producing structured summaries. It also does it slowly and expensively, at a volume that is starting to matter, or it cannot run where you need it to. Someone suggests fine-tuning a small model. Someone else says fine-tuning is a maintenance trap.
Both can be right. A small model adapted to one task can be fast, cheap and private. It can also be a model nobody knows how to retrain, built on data nobody had the right to use, that quietly degrades when inputs change. The question is whether a smaller model can meet your quality and latency requirements, and whether that is worth the added maintenance.
What the work involves
An evaluation before any training. I build a task evaluation with your reviewers: real inputs, expected outputs and a scoring method checked against human judgement. Everything that follows is measured on it.
An honest baseline. Before training anything, I measure the alternatives: the large model with careful prompting and examples, retrieval where the task depends on knowledge, and off-the-shelf small models. A good share of fine-tuning proposals end here, because a cheaper option already meets the bar. That is a useful result, not a failed project.
A rights-checked training plan. Each training source is checked with its owner: may it be used to train a model, does it contain personal data, and does any provider’s terms restrict using its outputs as training data? Data that fails the check is excluded, and the plan records why.
Train and compare. Candidate small models are fine-tuned on labelled examples or distilled from a larger teacher, with teacher outputs filtered and checked. Each is evaluated on held-out data for quality, latency and cost per task, beside the baseline.
Count the maintenance. Retraining frequency, drift monitoring, base-model upgrades and who owns them are part of the decision, not an afterthought.
The signature deliverable
You receive the baseline comparison, rights-checked training plan, task evaluation and deployment decision. Illustrative example of a baseline comparison:
| Approach | Task quality | p95 latency | Cost per 1,000 tasks (relative) | Maintenance |
|---|---|---|---|---|
| Large hosted model, tuned prompt | 94% | 2.8 s | 1.00 | Low |
| Small off-the-shelf model, same prompt | 78% | 0.4 s | 0.06 | Low |
| Small model fine-tuned on 1,500 labelled examples | 92% | 0.4 s | 0.07 | Retrain quarterly; drift monitor |
Illustrative example. Figures are placeholders showing the format, not results from a client.
How acceptance is judged
Acceptance thresholds are agreed after the baseline: the quality floor relative to the baseline, latency and cost targets, and the maintenance the owner is prepared to carry. A go requires the candidate to meet them on held-out data. A no-go is a complete, accepted outcome. On a go, deployment is accepted when the model meets the same thresholds in your environment and the owner has run the retraining pipeline once.
Ownership and handover
Models, training code, data lineage and the evaluation live in your accounts. The ML or engineering owner receives the retraining routine, drift monitoring and the criteria for deciding when to retrain or retire the model. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If the main question is where to run models, see private and local LLM deployment. If you want cheaper or faster inference without training anything, start with reduce LLM cost and latency. If you need a release gate for model changes generally, see production LLM evaluation. For broader ML pipelines and serving, see Python and machine learning engineering.