Industries · B2B SaaS

An AI feature that holds tenant boundaries and has a known cost per customer

If you are adding an AI feature to a multi-tenant B2B product, I scope and deliver it with tenant isolation designed into retrieval, prompts, caches and logs, a task-level evaluation your product team can rerun, and a per-customer cost model tied to your plans. Your team owns the feature, the evaluation and the maintenance plan at handover.

This is a good fit if…

  • You are adding an assistant, search, summarisation or agent feature to an existing multi-tenant product.
  • Enterprise prospects are asking how their data is separated, whether it is used for training and where it is processed.
  • Your finance or pricing lead cannot tell what the feature will cost per account once heavy users adopt it.
  • A prototype works on demo data, but nobody has tested it across tenants with very different data shapes and volumes.

Look elsewhere if…

  • You want an engineer to join your team under your own manager. Use the contract LLM engineer role instead.
  • The feature is live and the only problem is cost or latency. Use LLM cost and latency reduction instead.
  • Your customers need help deploying your AI product in their environments. Use the software vendor partner route.

What you get

Tenant-aware feature architecture, task evaluation and per-customer cost model

  • A tenant-aware architecture stating where tenant identity is enforced: retrieval, prompt assembly, caching, logging, evaluation data and any fine-tuning data.
  • A cross-tenant leak test suite that runs in CI and blocks a release on failure.
  • A task evaluation for each job the feature does, with results broken down by representative customer segment.
  • A per-customer cost model covering model calls, retrieval, storage and the usage distribution across your plans.
  • A maintenance plan for model and provider changes, evaluation re-runs and on-call ownership.

How it runs

  1. 01

    Architecture and tenancy review

    I map your tenancy model, data stores, identity and existing AI calls, then propose where isolation is enforced and what the feature must never do across accounts.

  2. 02

    Evaluation and cost baseline

    We define the feature's tasks and build an evaluation set across representative tenants, and I model cost per account from your real usage distribution.

  3. 03

    Build behind a flag

    The feature ships through your review process to internal accounts, then to a cohort of design-partner customers, with leak tests and evaluation in CI.

  4. 04

    Rollout plan and handover

    Results by segment, the cost model, the maintenance plan and the code are handed to your product and engineering owners.

What needs to be in place

  • A product owner with a defined feature and authority over its scope and plan placement.
  • An engineering owner whose team will maintain the feature.
  • A staging environment with multiple realistic tenants, or permission to build synthetic ones.
  • Your customer contract position on data use, sub-processors and data location, as agreed by your legal team.

Not included

  • Pricing or packaging decisions; the cost model informs them, your team makes them.
  • Drafting customer contracts, data processing agreements or security questionnaire answers.
  • Promised margin, adoption or retention figures.
  • Ongoing on-call for the feature unless agreed separately in writing.

The situation

Most B2B SaaS companies now have an AI feature on the roadmap, and many have a prototype. The prototype was built against one demo account, with a generous model and no limits. Three questions appear as soon as it heads towards real customers.

First, isolation. Your product already separates tenants in its database and permissions. An AI feature adds new places where data can cross accounts: a shared index, a prompt that includes cached context, a log that a support engineer reads, an evaluation set built from customer content.

Second, unit economics. A feature that costs very little per request can still cost more than the account pays when a handful of customers use it heavily. Pricing teams need a model, not an average.

Third, maintenance. The feature depends on a provider that will change models, prices and behaviour. Without an evaluation suite and an owner, quality drifts and nobody notices until a customer does.

What the work involves

Tenant-aware architecture. I trace every path the feature takes through your data and state where tenant identity is enforced. Retrieval filters are applied inside the query, not after it. Caches are keyed by tenant. Prompts never mix context from two accounts. Logs and traces that contain customer content inherit the same access rules as the product. Evaluation and any tuning data are tagged by tenant and kept within the boundaries your contracts allow. The rule is the one from my published work on agent safety: boundaries are enforced by the system, not requested of the model.

Cross-tenant leak tests. For each path, a test seeds two tenants with distinguishable content and asserts that neither can reach the other’s. These run in CI and block a release on failure.

Task evaluation. A SaaS AI feature usually does several jobs, such as summarising a record, answering a question over the customer’s documents or drafting an update. Each job gets its own evaluation set built with your product team, run across representative customer segments. Small and large tenants often behave differently, and an average hides that.

Per-customer cost model. Built from your real usage distribution, current provider prices and the feature’s call pattern, with the levers made explicit: caching, smaller models for simpler tasks, usage limits by plan.

I have integrated AI assistance into a live marketplace product, and I build agent infrastructure components through Neul Labs; both shaped this approach to product-grade AI features.

The signature deliverable, illustrated

Alongside the shipped feature, the cost model is often the artefact leadership reads first. Illustrative example with placeholder figures, not client data:

SegmentShare of accountsFeature tasks per account per monthRelative cost per accountLever
Small70%Low1 (baseline)None needed
Mid-market25%Moderate6Cache shared context
Enterprise5%High40Plan-level limits; smaller model for summaries

How acceptance is judged

The scope names the tasks, the evaluation thresholds per task, the leak tests that must pass and the cost assumptions to be validated. Your product owner accepts against those. A design-partner cohort tests the feature before general release, and their results are reported by segment.

Ownership and handover

Your engineering owner holds the code, the leak test suite, the evaluation harness and the maintenance plan; your product owner holds the cost model and rollout decisions. Engineers who will maintain the feature build it with me, so handover is a confirmation rather than a transfer.

When to choose something else

If you want a senior engineer working your own backlog, use the contract LLM engineer role. If the feature is live and too expensive or slow, see LLM cost and latency. For an evaluation gate on an existing feature, see production LLM evaluation. If your customers need hands-on help deploying your AI product, the software vendor partner route fits better. Other sectors are on the industries overview.

Questions buyers ask

Is a shared vector index safe for multiple tenants?

It can be, if tenant filters are applied inside the query rather than after retrieval, and if every write path tags documents correctly. Separate indexes or namespaces per tenant are simpler to reason about but cost more at scale. The architecture review chooses based on your tenant count, data sizes and enterprise commitments, and the leak test suite checks whichever option you pick.

How do you estimate cost per customer before launch?

From your own usage data. I take the distribution of activity across accounts, map it to the calls the feature makes per task, and price it at current provider rates with explicit assumptions. Heavy accounts usually dominate, so the model shows the cost at the median and the top percentiles, and where limits or caching change the picture.

Will customer data be used to train models?

Not unless your contracts allow it and you decide to. By default the design uses provider configurations that do not retain or train on your data, and keeps evaluation data inside your environment. Where a feature would benefit from tenant-specific tuning, that is an explicit product and legal decision, not an engineering default.

Can you work with our existing engineering team?

That is the normal pattern. I deliver the scoped outcome and pair with your engineers throughout, so the people who will maintain the feature build it with me. Priorities inside the scope stay with your product owner, and design decisions are written down as they are made, so the reasoning survives after the engagement ends.

What happens when the model provider changes their model?

The maintenance plan covers it: pinned model versions where the provider allows, the evaluation suite rerun before any switch, and a named owner for the decision. Without that, quality changes reach customers before anyone notices. The plan also lists which customer commitments, such as data location, constrain which replacement models you may use.

Related engagements

Industries · Regulated operations

Regulated operations

We need useful AI workflows in a controlled environment. How do we translate agreed policies into technical and review controls?

You get:Policy-to-workflow control mapping, evidence capture and pilot approval dependencies

Industries · Retail and marketplaces

Retail and marketplaces

Which search, merchandising and operations problems should we tackle first, and how will we measure whether the change helps?

You get:Retail-specific opportunity brief, evaluated retrieval or recommendation pilot and rollout plan

AI delivery · Evaluation

LLM evaluation

We want to change models without breaking customer workflows. Who can implement representative tests and release criteria?

You get:Engineering evaluation harness and remediation

AI delivery · Cost and latency

LLM cost and latency

Our AI feature is too slow or expensive. How do we identify improvements while checking quality and operational cost?

You get:Workload profile, quality-cost-latency frontier and measured optimisation backlog

Industries · Professional services

Professional-services firms

How should our firm sequence AI adoption across client confidentiality, partner review and billable delivery work?

You get:Firm-level use-case portfolio, client boundary model and governed pilot sequence

Industries · Education operations

Education operations

We need help with administrative research and communications, not automated student selection or grading. What workflow fits?

You get:Staff workflow pilot, source checks, restricted-data handling and human approval

Further reading

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics