AI delivery · Voice (selective)

Test a voice AI workflow's latency, consent and escalation before it meets callers

If you are planning a voice workflow that must respond quickly, hand callers to people cleanly and respect recording permissions, you can ask me to engineer and test it before wider deployment. I take voice work on selectively, after a fit check. Where it fits, I set a latency budget, map consent and retention dependencies, run scripted evaluation calls and design the human handoff.

This is a good fit if…

  • You have one defined voice workflow in mind, such as appointment handling, order status, call triage or after-hours intake, not a general-purpose phone agent.
  • Callers will notice delays, and you need to know where the latency goes before committing to a vendor or architecture.
  • Recording, transcription and retention must follow consent rules your legal or compliance owner has set.
  • Some callers must reach a person quickly and with context, and you need that handoff designed rather than improvised.

Look elsewhere if…

  • You want an always-on, managed voice service with support around the clock. I do not offer that.
  • The workflow is mostly text, documents or system updates. Use AI workflow automation or document workflows instead.
  • You need emergency, clinical, crisis or other safety-critical call handling. That needs specialist providers and is outside this work.
  • You want a vendor selected and contracted on your behalf. I can test candidates; procurement stays with you.

What you get

Latency budget, consent and retention dependencies, evaluation calls and handoff design

  • A latency budget for each stage of a turn (speech recognition, reasoning, tool calls, speech synthesis, network) measured on your workflow.
  • A map of consent, recording, transcription and retention dependencies, with the owner who must approve each.
  • A set of scripted evaluation calls covering normal, difficult and adversarial callers, with results recorded.
  • A handoff design: when the workflow transfers to a person, what context travels with the call, and what happens out of hours.
  • A go, revise or stop view on wider deployment, with the evidence behind it.

How it runs

  1. 01

    Brief and fit check

    You describe the workflow, callers, systems and constraints. I check whether the work fits my capability and current capacity and tell you plainly if a voice specialist would serve you better.

  2. 02

    Latency budget and dependency map

    I break down a conversational turn, measure each stage on candidate components and map the consent, recording and retention decisions your owners must make.

  3. 03

    Build a test workflow

    A bounded version of the workflow is built against test lines and test systems, with escalation paths and logging.

  4. 04

    Evaluation calls

    Scripted calls cover accents, background noise, interruptions, silence, off-script requests and attempts to manipulate the agent, measured against agreed criteria.

  5. 05

    Decision and handover

    Results, the handoff design and the operating dependencies are handed to the product or customer operations owner, with a go, revise or stop recommendation.

What needs to be in place

  • One defined workflow with a product or customer operations owner.
  • A position from your legal or compliance owner on call recording, transcription and retention, or agreement to obtain one.
  • Test telephony access and test versions of any systems the workflow reads or updates.
  • People who can act as callers and reviewers for the evaluation calls.

Not included

  • An always-on or round-the-clock managed voice service.
  • Legal advice on consent, recording or data-protection obligations.
  • Emergency, clinical or other safety-critical call handling.
  • Telephony carrier contracts, number porting or contact-centre platform replacement.
  • Voice cloning or impersonation of real people.

Taken on selectively

Voice workflows are a part of AI delivery I take on selectively, after a fit check against both capability and current capacity. This page does not describe an always-on service. If your requirement is better served by a dedicated voice vendor or specialist, I will tell you that at the first conversation rather than stretch to fit.

The situation

A voice workflow looks simple in a demo: someone calls, a model answers, a task gets done. On real calls the problems are different from text. Every pause is audible. Callers interrupt, mumble, call from cars and change their minds. Some must reach a person immediately, and they should not have to repeat themselves. Every call raises questions about recording, transcription, storage and who may listen later.

Most of these issues surface only after a vendor is chosen and the workflow is live. This page is for testing them first: what should you measure before wider deployment, and what must be decided by people other than the engineer?

What the work involves

A latency budget per turn. A conversational turn passes through speech recognition, end-of-speech detection, the model, any tool calls, speech synthesis and the network, in both directions. I measure each stage on your workflow and candidate components, at typical and slow percentiles, and set a budget per stage. Tool calls to slow internal systems are often the hidden cost; they may need caching, prefetching or a holding phrase.

Consent and retention dependencies. I map every point where audio or transcripts are recorded, processed, stored or sent to a third party, and list the decisions your legal or compliance owner must make: when callers are told, how opt-outs work, retention periods, and which providers may process recordings. The design follows those decisions; it does not make them.

Handoff by design. We define the conditions for transferring to a person (explicit request, repeated misunderstanding, sensitive topics, high-value actions), what context travels with the transfer, and what happens out of hours. A handoff that drops the context is a failure even if the transfer works.

Evaluation calls, not demos. Scripted calls cover the realistic range: accents, noise, interruptions, silence, off-script requests, callers trying to manipulate the agent into actions it should not take. Each call is scored against agreed criteria and recorded under the agreed consent rules.

The signature deliverable

You receive the latency budget, consent and retention dependencies, evaluation calls and handoff design. Illustrative example of a latency budget:

StageBudget (p90)Measured (p90)Note
End-of-speech detection300 ms420 msTuning needed for noisy lines
Speech recognition (final)200 ms180 msWithin budget
Model response, first token400 ms650 msShorter context; smaller model for routing
Order lookup tool call300 ms1,100 msPrefetch on caller identification
Speech synthesis, first audio200 ms190 msWithin budget

Illustrative example. Figures are placeholders showing the format, not results from a client.

How acceptance is judged

Before evaluation calls start, the owner and I agree criteria: the latency target at stated percentiles, task completion on scripted calls, correct escalation on every call that should escalate, no prohibited action taken, and consent behaviour matching the documented rules. The result is a go, revise or stop recommendation, and stop is a legitimate outcome.

Ownership and handover

The test workflow, call scripts, results and handoff design belong to you. The product or customer operations owner holds the decision on wider deployment and the operating dependencies. Ongoing operation stays with your team or provider. I do the work personally; any specialist help, for example in telephony, is disclosed and approved by you first.

Boundaries

If the workflow does not need to happen on calls, AI workflow automation or document workflows will be simpler. If the voice channel is part of a support operation that needs knowledge and escalation practices, see AI enablement for customer support. If your data sources are not ready for any assistant to use, start with AI data readiness.

Questions buyers ask

Why is voice work taken on selectively?

Voice adds constraints that text workflows do not have: real-time latency, telephony, consent and callers who cannot see what the system is doing. I take it on when the workflow is well defined and my capacity allows the attention it needs. If a dedicated voice specialist or vendor would serve you better, I will say so at the fit check.

How fast does a voice agent need to be?

Fast enough that callers do not talk over it or assume the line has dropped. The acceptable delay depends on the workflow and callers, so we agree a target per turn and measure the real distribution, including slow tail cases, rather than quoting a single typical figure.

Do we need a particular voice platform?

No. The latency budget and evaluation calls are designed to compare candidate components, whether a managed voice platform or separate speech and model services. The choice depends on latency, cost, data handling and what your team can operate. I do not resell any platform.

What about recording consent?

Your legal or compliance owner decides the rules; I map where the workflow records, transcribes, stores or sends audio and text, and make sure the design can follow those rules, including telling callers and honouring opt-outs. I do not provide legal advice on what the rules should be.

What happens after the evaluation?

You get a go, revise or stop recommendation with the evidence. If you proceed, wider deployment is a separate, scoped decision. Ongoing operation of a voice workflow is not something I provide as an always-on service; your team or a provider runs it.

Describe what needs to work

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics