Contract engineering · LLM applications

A senior LLM engineer shipping your funded backlog in your codebase

If you have a funded LLM backlog and an application team that owns it, you can contract me as an LLM engineer. I join under your engineering manager, agree a written remit and access list in the first days, then implement model integrations, structured outputs and evaluation in your codebase. You end with shipped features, an evaluation suite and a handover plan.

This is a good fit if…

  • You have a funded backlog of LLM features in an existing product, and an engineering team that owns the application.
  • Prototypes work in a notebook, but nobody has made them reliable enough to ship: malformed outputs, drifting quality, unpredictable cost.
  • Your engineers are capable but have not built evaluation, structured output handling or provider fallbacks before.
  • You want one senior engineer under your manager, not a separate project team building beside yours.

Look elsewhere if…

  • Your main problem is retrieval quality or permission-aware search. Brief the contract RAG engineer instead.
  • The work is mainly agents that take actions in other systems. The AI agent engineer role fits that better.
  • You want an evaluation gate designed and delivered as a fixed outcome. Commission production LLM evaluation instead.

What you get

Named contractor remit, delivery backlog, access prerequisites and handover plan

  • LLM features shipped through your normal review and release process, not maintained on a side branch.
  • An evaluation set for each feature, run in CI, so prompt or model changes show their effect before release.
  • Structured outputs validated against schemas, with retries and safe failure paths instead of silent bad data.
  • Per-call latency and cost visible in your existing monitoring.
  • A handover plan your team has already started using before the contract ends.

Responsibilities I can own

  • Agree a written remit with your manager: what I own, what I advise on, and what stays with the team.
  • Implement model calls behind a thin provider interface, with timeouts, retries and fallbacks.
  • Enforce structured outputs with schema validation and explicit handling of refusals and malformed responses.
  • Build per-feature evaluation sets with your product owners and wire them into CI.
  • Put prompts and model versions under version control, with change notes in pull requests.
  • Add logging that captures what is needed to debug quality without retaining personal data unnecessarily.
  • Pair with your engineers so each pattern is understood and repeatable.

Stack fit

  • Python
  • TypeScript
  • LLM provider APIs
  • Open-weight models
  • JSON Schema / Pydantic / Zod
  • Evaluation harnesses
  • PostgreSQL
  • OpenTelemetry
  • CI pipelines
  • Docker

Onboarding I need from you

  • Repository, staging and secrets access through your normal joiner process.
  • A named engineering manager who owns the backlog and accepts work.
  • API access to the model providers your organisation has approved, with usage limits agreed.
  • A product owner or domain expert who can judge output quality for an hour or two a week.
  • Your data-handling rules for prompts and logs, in writing, before production traffic is used.

Reporting

I report to your engineering manager, work your sprint cadence, and send a short written update each week covering shipped changes, evaluation results and open risks.

How it runs

  1. 01

    Brief and fit check

    You send the backlog or role brief. I reply with questions, a plain view on fit, and an availability window checked against your start date.

  2. 02

    Remit and access in the first days

    We write down what I own, what I advise on, the access I need and who approves it. Blocked access is the most common reason contracts stall, so this comes first.

  3. 03

    Work the backlog

    Features ship through your review process, each with an evaluation run attached. Your manager sets the order and can change it.

  4. 04

    Handover

    A handover plan agreed mid-contract, then executed: evaluation suite, runbooks, remaining backlog and known risks.

What needs to be in place

  • An application your team owns and deploys, with a backlog of LLM work that is funded and prioritised.
  • An engineering manager with authority over that backlog.
  • Approved model providers or hosting, and a written basis for what data may be sent to them.
  • A contract route agreed up front: direct, via your agency or a partner, and your IR35 or equivalent determination.

Not included

  • Ownership of product direction or line management of your engineers.
  • Promised accuracy or quality figures. Evaluation shows what improved and what did not.
  • Model training from scratch. Small-model adaptation is a separate scoped service.
  • Out-of-hours on-call support unless agreed separately in writing.
  • Substitution by another engineer. Any specialist help is named and approved by you first.

When an LLM contract is the right buy

The usual starting point is a product team that has already proved an LLM feature can work. There is a demo, perhaps an internal beta, and a funded backlog: summarise this record, extract fields from that document, draft a reply, classify incoming requests. What there is not is someone who has taken this kind of feature to production before. The prototype returns malformed JSON one time in fifty, quality shifts when the provider updates a model, nobody can say whether last week’s prompt change helped, and the cost per request is a guess.

You do not need a separate project team for this. You need a senior engineer who has done it before, sitting in your team, implementing the backlog in your codebase and leaving the patterns behind.

What the work involves

Most LLM backlogs share a handful of engineering problems underneath the feature names. I work through them in the order your manager sets, but they tend to look like this:

  • Output you can trust structurally. Every model response that feeds code is validated against a schema. Refusals, truncation and malformed output have explicit paths: retry, repair, fall back or fail visibly. Nothing writes bad data quietly.
  • Quality you can measure. Each feature gets a small evaluation set built with the people who know what a good answer looks like. It runs in CI, so a prompt edit or model change shows its effect in the pull request.
  • Change you can trace. Prompts and model identifiers live in version control. A quality regression can be traced to a specific change rather than argued about.
  • Cost and latency you can see. Tokens, latency and failure rates per feature go into your existing monitoring, with budgets your team agrees.
  • Data you can account for. What is sent to which provider, what is logged and for how long, written down and enforced in code.

The signature deliverable

The contract starts and ends with documents your team keeps: a named contractor remit, a delivery backlog, the access prerequisites and a handover plan. The remit matters most, because it stops the role drifting. Illustrative example of a remit extract:

AreaI ownI adviseTeam owns
Ticket-summary featureImplementation, evaluation set, schema validationPrompt wording with support leadsRelease decision
Provider interfaceDesign and first implementationChoice of fallback modelContract with the provider
Logging and redactionImplementation to your policyRetention periodThe policy itself
CI evaluation gateHarness and thresholds proposalThreshold valuesEnforcing the gate

Illustrative example showing the format, not a record from a client engagement.

The access prerequisites list sits alongside it: each system, the level of access, who approves it, and the date it is needed. Contracts most often lose their first weeks to access, so this list is agreed before day one where possible.

How acceptance is judged

Your engineering manager accepts each change through normal review. Every LLM feature change carries an evaluation run in the pull request, compared against the previous baseline, so acceptance is based on evidence rather than a reviewer reading a few outputs. At the end, the measurable output is the shipped features, the evaluation suite your team can run, and a backlog with what is left in priority order.

Ownership and handover

Everything is built in your repositories, under your conventions, reviewed by your engineers. I agree the handover plan with your manager around the midpoint, naming who will own each area, and then work towards it: those engineers pair on the relevant changes and run the evaluation suite themselves before I leave. The written handover covers how to add an evaluation case, how to change a prompt or model safely, the provider interface, the remaining backlog and known risks.

When to choose something else

If the hard part is retrieval, the RAG engineer contract is a better fit. If the features take actions in other systems, see the AI agent engineer role. If you want a quality gate delivered as a fixed outcome, commission production LLM evaluation. If you are moving between providers, LLM provider migration is scoped for exactly that. For a smaller recurring load, consider part-time AI engineering capacity.

Questions buyers ask

What evidence should we ask an LLM contractor for?

A current CV that says what the person personally built, work you can inspect, and a design conversation about one of your backlog items. Ask how they would handle a malformed model response, how they would know a prompt change made things worse, and what they log. Vague answers about prompt engineering are a warning sign; specific answers about schemas, evaluation sets and failure paths are not.

Which model providers do you work with?

Whichever your organisation has approved. I keep provider calls behind a thin interface so a change of model or vendor is a contained piece of work rather than a rewrite. I am independent and have no commercial arrangement with any model provider, so the recommendation is about fit for your workload and data rules.

How is this different from commissioning an LLM evaluation project?

A contract buys senior capacity under your manager across the whole backlog; evaluation is one part of how I work. A commissioned evaluation project is a fixed outcome, such as a production quality gate, that I am accountable for delivering against acceptance criteria. If evaluation is the whole requirement, the scoped service is the cleaner buy.

What contract basis do you work on?

Directly, through your preferred agency, or through a delivery partner, at the published day rate. Your organisation makes the IR35 or equivalent status determination. I will explain the working practices so your assessment reflects them, but I do not give tax advice.

What if our backlog changes halfway through?

That is normal and the reason the remit is written separately from the backlog. Your manager reorders work as priorities shift. If the change moves the work outside the remit, for example from LLM features to platform operations, we update the remit in writing so both sides know what I own.

Related engagements

Contract engineering · Agents

AI agent engineer

We need someone to build tool-using workflows with reliable state, permissions and recovery, not another agent demo. What skills and scope fit?

You get:Contract brief covering tool actions, state, evaluation and production ownership

Contract engineering · Retrieval

RAG engineer

We need embedded capacity to improve retrieval, permissions and answer evaluation in an existing team. What should the contract deliver?

You get:Retrieval backlog, relevance evaluation plan, access-control tasks and handover

AI delivery · Evaluation

LLM evaluation

We want to change models without breaking customer workflows. Who can implement representative tests and release criteria?

You get:Engineering evaluation harness and remediation

Contract engineering · Part-time

Part-time AI engineer

We have meaningful work but not a full-time workload. How can a part-time contract avoid fragmented ownership and waiting time?

You get:Agreed weekly capacity, work-in-progress limit, communication window and decision SLA

Contract engineering · Enablement

AI enablement engineer

We need a hands-on engineer who can implement workflows and coach users rather than sell a training-only programme. How should the role be written?

You get:Embedded implementation-and-coaching remit, workflow backlog and adoption handover

Contract engineering · Customer deployment

Forward-deployed engineer

Our customer deployments need an engineer who can work across the product team and client environment. What responsibilities should the contract include?

You get:Customer deployment remit, integration dependencies and acceptance ownership

Further reading

Send an engineering brief

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics