Taken on selectively
Voice workflows are a part of AI delivery I take on selectively, after a fit check against both capability and current capacity. This page does not describe an always-on service. If your requirement is better served by a dedicated voice vendor or specialist, I will tell you that at the first conversation rather than stretch to fit.
The situation
A voice workflow looks simple in a demo: someone calls, a model answers, a task gets done. On real calls the problems are different from text. Every pause is audible. Callers interrupt, mumble, call from cars and change their minds. Some must reach a person immediately, and they should not have to repeat themselves. Every call raises questions about recording, transcription, storage and who may listen later.
Most of these issues surface only after a vendor is chosen and the workflow is live. This page is for testing them first: what should you measure before wider deployment, and what must be decided by people other than the engineer?
What the work involves
A latency budget per turn. A conversational turn passes through speech recognition, end-of-speech detection, the model, any tool calls, speech synthesis and the network, in both directions. I measure each stage on your workflow and candidate components, at typical and slow percentiles, and set a budget per stage. Tool calls to slow internal systems are often the hidden cost; they may need caching, prefetching or a holding phrase.
Consent and retention dependencies. I map every point where audio or transcripts are recorded, processed, stored or sent to a third party, and list the decisions your legal or compliance owner must make: when callers are told, how opt-outs work, retention periods, and which providers may process recordings. The design follows those decisions; it does not make them.
Handoff by design. We define the conditions for transferring to a person (explicit request, repeated misunderstanding, sensitive topics, high-value actions), what context travels with the transfer, and what happens out of hours. A handoff that drops the context is a failure even if the transfer works.
Evaluation calls, not demos. Scripted calls cover the realistic range: accents, noise, interruptions, silence, off-script requests, callers trying to manipulate the agent into actions it should not take. Each call is scored against agreed criteria and recorded under the agreed consent rules.
The signature deliverable
You receive the latency budget, consent and retention dependencies, evaluation calls and handoff design. Illustrative example of a latency budget:
| Stage | Budget (p90) | Measured (p90) | Note |
|---|---|---|---|
| End-of-speech detection | 300 ms | 420 ms | Tuning needed for noisy lines |
| Speech recognition (final) | 200 ms | 180 ms | Within budget |
| Model response, first token | 400 ms | 650 ms | Shorter context; smaller model for routing |
| Order lookup tool call | 300 ms | 1,100 ms | Prefetch on caller identification |
| Speech synthesis, first audio | 200 ms | 190 ms | Within budget |
Illustrative example. Figures are placeholders showing the format, not results from a client.
How acceptance is judged
Before evaluation calls start, the owner and I agree criteria: the latency target at stated percentiles, task completion on scripted calls, correct escalation on every call that should escalate, no prohibited action taken, and consent behaviour matching the documented rules. The result is a go, revise or stop recommendation, and stop is a legitimate outcome.
Ownership and handover
The test workflow, call scripts, results and handoff design belong to you. The product or customer operations owner holds the decision on wider deployment and the operating dependencies. Ongoing operation stays with your team or provider. I do the work personally; any specialist help, for example in telephony, is disclosed and approved by you first.
Boundaries
If the workflow does not need to happen on calls, AI workflow automation or document workflows will be simpler. If the voice channel is part of a support operation that needs knowledge and escalation practices, see AI enablement for customer support. If your data sources are not ready for any assistant to use, start with AI data readiness.