The risk that is specific to product teams
Product discovery produces a lot of raw material: interview transcripts, support tickets, sales call notes, survey free text, feature requests. Turning it into themes and requirements used to take days of reading and sticky notes. AI assistants can now produce a tidy synthesis in minutes, and most product managers have tried it.
The danger is subtle. A model asked “what are customers’ main frustrations?” will give a fluent, plausible answer whether or not the transcripts support it. It may blend three customers into a theme, overweight one memorable quote, or add a frustration that sounds right but nobody said. That answer then appears in a synthesis deck, then in a requirements document, and becomes “what customers told us”. Product teams end up building on generated opinion presented as evidence.
What the work involves
Source-linked synthesis. The model helps sort and cluster raw material, but every theme in the synthesis must link to the specific quotes, tickets or responses that support it, with a count and a denominator (“7 of 18 interviews”, not “many customers”). Themes without links are removed or reclassified.
Evidence labels. The team agrees four labels and uses them everywhere: observed (seen in usage data or a session), reported (said by a customer, with a source), inferred (the team’s interpretation of reported or observed evidence), and generated hypothesis (suggested by a model, untested). The labels make the strength of each claim visible at a glance.
Reviewed requirements. Each requirement cites the themes and labels it rests on. Requirements built mainly on inferences or generated hypotheses are marked for validation before build. A second reviewer, usually another product manager or a researcher, checks labels and links on a sample.
I led a product engineering transformation covering teams, process and stack at a peer-to-peer marketplace, where getting from customer evidence to engineering work cleanly was a large part of the job.
The signature deliverable
You end with source-linked discovery synthesis, evidence labels and a reviewed requirements workflow. Illustrative example of a synthesis extract:
| Theme | Support | Label | Sources |
|---|---|---|---|
| Admins struggle to bulk-invite users | 7 of 18 interviews; 23 tickets in quarter | Reported | Interview IDs, ticket query |
| Invite flow abandoned at role selection | Funnel drop at step 3 | Observed | Analytics dashboard link |
| Admins want role templates | Team interpretation of the two themes above | Inferred | Links to themes |
| Admins would pay for SSO provisioning | Model suggestion; no customer said this | Generated hypothesis | To test in next round |
Illustrative example. Rows show the format, not a client’s research.
How acceptance is judged
Before the pilot, we trace a sample of themes and requirements from recent work back to their sources and record how many can be traced. During the pilot we measure the same traceability, along with time from end of interviews to reviewed synthesis and how often the second reviewer disagrees with a label. The head of product accepts the pilot. The adoption measure is whether product managers use the labels in planning discussions and whether requirements arrive with evidence attached.
Ownership and handover
The product operations owner holds the templates, labels and review routine, and coaches new product managers on them. Your research lead, if you have one, owns the consent and data rules. The handover note covers how to add a source type and how to run the label review.
Boundaries
This pilot is for product discovery and requirements. If engineering teams are adopting coding assistants, see AI engineering enablement for software teams. If marketing is preparing research and content for publication, see AI enablement for marketing teams. For a broader view across business teams, see practical AI enablement for business teams.