The situation
Most B2B SaaS companies now have an AI feature on the roadmap, and many have a prototype. The prototype was built against one demo account, with a generous model and no limits. Three questions appear as soon as it heads towards real customers.
First, isolation. Your product already separates tenants in its database and permissions. An AI feature adds new places where data can cross accounts: a shared index, a prompt that includes cached context, a log that a support engineer reads, an evaluation set built from customer content.
Second, unit economics. A feature that costs very little per request can still cost more than the account pays when a handful of customers use it heavily. Pricing teams need a model, not an average.
Third, maintenance. The feature depends on a provider that will change models, prices and behaviour. Without an evaluation suite and an owner, quality drifts and nobody notices until a customer does.
What the work involves
Tenant-aware architecture. I trace every path the feature takes through your data and state where tenant identity is enforced. Retrieval filters are applied inside the query, not after it. Caches are keyed by tenant. Prompts never mix context from two accounts. Logs and traces that contain customer content inherit the same access rules as the product. Evaluation and any tuning data are tagged by tenant and kept within the boundaries your contracts allow. The rule is the one from my published work on agent safety: boundaries are enforced by the system, not requested of the model.
Cross-tenant leak tests. For each path, a test seeds two tenants with distinguishable content and asserts that neither can reach the other’s. These run in CI and block a release on failure.
Task evaluation. A SaaS AI feature usually does several jobs, such as summarising a record, answering a question over the customer’s documents or drafting an update. Each job gets its own evaluation set built with your product team, run across representative customer segments. Small and large tenants often behave differently, and an average hides that.
Per-customer cost model. Built from your real usage distribution, current provider prices and the feature’s call pattern, with the levers made explicit: caching, smaller models for simpler tasks, usage limits by plan.
I have integrated AI assistance into a live marketplace product, and I build agent infrastructure components through Neul Labs; both shaped this approach to product-grade AI features.
The signature deliverable, illustrated
Alongside the shipped feature, the cost model is often the artefact leadership reads first. Illustrative example with placeholder figures, not client data:
| Segment | Share of accounts | Feature tasks per account per month | Relative cost per account | Lever |
|---|---|---|---|---|
| Small | 70% | Low | 1 (baseline) | None needed |
| Mid-market | 25% | Moderate | 6 | Cache shared context |
| Enterprise | 5% | High | 40 | Plan-level limits; smaller model for summaries |
How acceptance is judged
The scope names the tasks, the evaluation thresholds per task, the leak tests that must pass and the cost assumptions to be validated. Your product owner accepts against those. A design-partner cohort tests the feature before general release, and their results are reported by segment.
Ownership and handover
Your engineering owner holds the code, the leak test suite, the evaluation harness and the maintenance plan; your product owner holds the cost model and rollout decisions. Engineers who will maintain the feature build it with me, so handover is a confirmation rather than a transfer.
When to choose something else
If you want a senior engineer working your own backlog, use the contract LLM engineer role. If the feature is live and too expensive or slow, see LLM cost and latency. For an evaluation gate on an existing feature, see production LLM evaluation. If your customers need hands-on help deploying your AI product, the software vendor partner route fits better. Other sectors are on the industries overview.