← Blog

AI voice agents: where they work, and where they do not

AI voice agents work on bounded, structured calls: out-of-hours and overflow enquiries, bookings, status checks and routing. They struggle with complex, emotional or high-stakes calls, mishear some accents, and can transcribe words nobody said. Design for that: tell callers it is automated, read back critical details, escalate by rule, and test before real calls.

AI agentsZegaware Engineering14 min read

Last updated: 9 October 2026

A voice agent is the most exposed AI system a business can run. It speaks for you, in real time, to a member of the public, with no chance to edit the answer before it is heard. When it works, a caller at eleven at night gets an appointment booked instead of a voicemail. When it does not, it mishears a name, invents a detail, or keeps a distressed caller talking to software that cannot help them.

This guide is part of our AI agents coverage and builds on how to build a reliable AI agent. It covers where voice agents earn their place, where they do not, and the engineering and legal facts most vendor pages skip, including one widely repeated "rule" that does not exist.

What an AI voice agent is

An AI voice agent is software that holds a spoken conversation, usually over a phone line, and takes actions during it: looking up an order, booking an appointment, routing the call. There are three common designs. OpenAI's developer documentation describes a chained voice pipeline of separate speech-to-text, language-model and text-to-speech stages; a realtime speech-to-speech model that is used to "interpret audio, decide what to do, and respond in speech" in one session; and full-duplex models that listen and speak at the same time [1].

The choice is a trade-off, not a ranking. The chained design is slower, but every stage is visible. OpenAI's own example is that you "might store the transcript, run policy checks before the text agent responds, call internal systems, then generate speech only after the workflow reaches an approved answer" [1]. A speech-to-speech model handles the rhythm of conversation more naturally, but there is less to inspect between what the caller said and what the agent did. For a business that may one day need to prove what its agent said and why, that inspectability is worth a great deal.

Where voice agents work

Voice agents do well on calls that are bounded, frequent and structured: the caller wants one of a small number of things, the information needed is well defined, and the outcome can be checked. Out-of-hours and overflow calls, qualifying an enquiry, booking into a calendar, checking an order status and routing to the right team all fit that description.

The independent evidence points the same way. A study of an LLM-based telephone survey system, run on real calls with 2,739 respondents in Peru, found that "overall data quality approached human-led standards for structured items", while the system's "probing for qualitative depth was more limited than human interviewers" [2]. Structured questions with clear answers went well. The open-ended, human part of the conversation did not.

That matches how we design them. A voice agent should own a defined set of call types, collect specific facts, write them to a system you can check, and hand everything else to a person.

Where they do not

Voice agents struggle where a call needs judgement, authority or empathy, or where the cost of mishearing is high. Complaints, distressed or vulnerable callers, anything that resembles regulated advice, disputes about money, and calls where the caller does not yet know what they need are all poor fits.

Speech recognition is also less even than demos suggest. A study published in PNAS tested five commercial speech-recognition systems and found "an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers" [3]. That is US data from systems of the time, but the lesson travels: accuracy depends on who is speaking. A 2025 study of Newcastle English found that recognition errors "directly correlate with regional dialectal features" [4]. A caller with a strong regional accent should not get a worse service than one without, and you will not know whether they do unless you test for it.

Benchmarks of complete voice agents tell the same story. EVA-Bench, published in May 2026, evaluated 12 systems across three architectures and found that "no system simultaneously exceeds 0.5" on both its accuracy and experience measures, and that "accent and noise perturbations expose substantial robustness gaps" [5].

Finally, do not use a voice agent to verify identity by voice. The Home Office's latest fraud assessment says it is "almost certain that criminals will increasingly adopt generative artificial intelligence (GenAI) technology such as deepfakes, Large Language Models (LLMs), and voice cloning to enable fraud" [6], and the ICO warns that "AI-generated audio and video can be used to impersonate colleagues or IT staff" [7]. A voice that sounds right is no longer evidence of who is speaking.

The latency budget

Human conversation is fast. A study of turn-taking across ten languages found that responses typically land between 0 and 200 milliseconds after the previous speaker stops, with a mean gap of "+208 ms" across the full dataset [8].

Many vendor pages cite a 150-millisecond target from ITU-T Recommendation G.114. That is a misreading. G.114 is about one-way transmission time on the network between two people: it says that below 150 milliseconds most applications "will experience essentially transparent interactivity", and that "a one-way delay of 400 ms should not be exceeded for general network planning" [9]. It describes the phone line, not how quickly software should think.

Realistic figures are much slower. Twilio, which sells the telephony many voice agents run on, published median benchmarks in November 2025 for a chained agent: a mouth-to-ear turn gap of 1,115 milliseconds, with an upper limit of 1,400, made up largely of speech-to-text (350), the language model's time to first token (375) and text-to-speech (100) [10]. Twilio describes these as "starting benchmarks, not best-in-class" [10]. Even a well-built agent replies around five times more slowly than a person does.

So design for the delay rather than against it. Keep replies short. Acknowledge a request before a slow lookup instead of leaving dead air. Avoid stacking tool calls into a single turn: as we covered in why AI agents fail in production, latency grows with every round trip. And measure the median and the 95th percentile, not the average [1], because the call a customer remembers is the slow one.

Interruptions, corrections and real phone lines

Real callers do not wait their turn. They interrupt, say "mm-hm" while the agent is talking, pause mid-sentence and change their minds. Researchers have built benchmarks specifically for these behaviours: Full-Duplex-Bench evaluates "pause handling, backchanneling, turn-taking, and interruption management" [11]. Its 2026 successor tested six systems on real human recordings with natural disfluencies and found that "self-correction handling and multi-step reasoning under hard scenarios remain the most consistent failure modes" across all of them [12]. The caller who says "Tuesday, no, sorry, Wednesday" is exactly the caller a voice agent is most likely to book on the wrong day.

The trade-offs in that benchmark were sharp. The chained pipeline never failed to take its turn but was the slowest to complete tasks, while the fastest system had the lowest turn-taking rate [12]. The authors conclude that "the next frontier for voice agents is not just reducing latency" [12].

The phone line itself makes all of this harder. Audio from a standard phone call is narrowband, "typically an 8khz sample rate" in Twilio's words, and Twilio recommends a speech model built for phone audio, with touch-tone keypad input "as an alternative input method when speech recognition fails" [13]. A demo recorded on a laptop microphone in a quiet room tells you very little about a caller on a mobile in a car park.

When the agent mishears, or makes something up

Speech recognition can do worse than mishear. A study presented at ACM FAccT 2024 found that "roughly 1% of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio", and that 38% of those hallucinations included explicit harms [14]. That study examined one model in 2023, so the rate is not a current figure for every system, but the failure mode is real: the transcript an agent acts on may contain words the caller never said.

The business owns the consequences. In Moffatt v Air Canada, a Canadian tribunal rejected the airline's argument that its chatbot was in effect "a separate legal entity that is responsible for its own actions", calling it "a remarkable submission", and held that "It makes no difference whether the information comes from a static page or a chatbot" [15]. It was a small-claims decision about a text chatbot in British Columbia and does not bind a UK court, but there is no reason to expect a UK regulator or court to take a more generous view.

The engineering response is to assume errors will happen and contain them:

  • Read critical details back. Names, dates, amounts and addresses are confirmed with the caller before anything is written to a system.
  • Route the dangerous cases by rule. An emergency keyword should trigger a fixed escalation, not depend on the model's judgement in the moment.
  • Scope the agent's authority. OWASP names "excessive functionality", "excessive permissions" and "excessive autonomy" as the root causes of excessive agency, and recommends human-in-the-loop control "to require a human to approve high-impact actions before they are taken" [16]. An agent that can book an appointment does not need to be able to cancel an account.
  • Treat what the caller says as untrusted input. Spoken input can carry prompt injection just as typed input can, so authorisation belongs in your systems, not in the model's instructions [16].
  • Keep a record you can replay. Every call should be transcribed, logged and reviewable, so you can show what the agent heard, decided and did. Our guide to AI agent audit trails covers what that record should contain.

What the law says in the UK and EU

Vendor content is weakest here, so it is worth separating law, guidance and good practice. This is not legal advice: confirm the position for your own use case with a solicitor.

There is no "five-second" disclosure rule. A claim circulates on vendor blogs that Ofcom requires AI voice agents to identify themselves within five seconds of answering. We found no Ofcom, ICO or government source for it. What Ofcom has said, in research published on 25 September 2026, is that "It should be very clear to telecoms customers when they are interacting with a GenAI application", and that firms should ensure "human support remains available when they need it" [17]. That is an expectation of telecoms firms, with no time limit attached.

Outbound automated marketing calls are tightly restricted. Regulation 19 of the Privacy and Electronic Communications Regulations (PECR) prohibits transmitting "communications comprising recorded matter for direct marketing purposes by means of an automated calling or communication system" without the recipient's prior consent [18]. The ICO's guidance is explicit that general marketing consent, or consent to live calls, does not cover automated calls [19]. That guidance does not mention AI, and whether a live, generated voice counts as "recorded matter" has not been settled. Our reading is that the question is open, and that anyone planning outbound AI calls should assume regulation 19 applies until the ICO says otherwise. The ICO's September 2025 fines of £250,000 and £300,000, often described online as "AI voice" cases, in fact concerned "scripted lines recorded by voice actors and played by call agents abroad" [20]. Calls that customers make to you are not direct marketing, so regulation 19 does not reach them.

Callers in the EU must be told they are talking to an AI. Article 50(1) of the EU AI Act requires providers to ensure that AI systems "intended to interact directly with natural persons" are designed so that people "are informed that they are interacting with an AI system, unless this is obvious" [21]. The rule applies from 2 August 2026 [22], and the Act reaches providers outside the EU "where the output produced by the AI system is used in the Union" [21]. The European Commission's guidelines name voice assistants explicitly and expect, in "voice-based or telephony contexts, explicit spoken statements at the beginning of the interaction", with periodic reminders in longer calls; distinct audio tones alone "are not considered sufficient by themselves" [23]. The guidelines also list vocal tone, "robot voice vs. genuine human-sounding voice", among the factors that decide whether the AI is obvious [23]. They are non-binding, but they are the clearest statement yet of what regulators expect.

UK data protection applies to every call. If you record or transcribe a call, UK GDPR requires you to tell the caller, and the ICO is clear that you "must provide privacy information to individuals at the time you collect their personal data from them" [24]. Understanding what a caller says is not, by itself, special category data: the ICO says biometric data "only becomes this if you use it to uniquely identify someone" [25], which is one more reason to avoid voiceprint authentication. And where an agent takes a decision with "a legal effect" or "a similarly significant effect" on the caller, with "no meaningful human involvement", the safeguards in the UK GDPR's Articles 22A to 22D apply, including the right to "obtain human intervention" [26].

Our position is simple. Every voice agent we build tells the caller, in words, at the start of the call, that it is an automated assistant. In the UK that is good practice rather than an explicit statutory duty; for callers in the EU it is the law. Either way, a business that lets callers believe they are speaking to a person is taking a risk it has no need to take.

Handing over to a human

Escalation should be a designed behaviour, not a fallback for when the agent gets confused. Agree the rules before go-live: which call types always go to a person, which words or situations trigger an immediate transfer, what happens out of hours, and what context the person receives so the caller does not have to start again. A good handover passes on the transcript and the facts already collected. A bad one drops the caller into a queue with nothing. People still value human support when they need it [17], and the handover is where that is won or lost.

Testing before real calls

A voice agent should hear thousands of calls before it hears a customer. OpenAI's guidance gives a sensible progression: start with synthetic speech, then "Replay representative human recordings", then "Use an independent simulated caller for continuous, multi-turn conversations" [1]. It recommends testing "input recognition across accents, background noise, language switches, names, and numbers", and checking the outcome, not just the conversation: "For a booking assistant, listen to the confirmation and check that the correct appointment was saved" [1].

In practice that means a test set built from your own call types, deliberately difficult callers (interruptions, corrections, strong accents, noisy lines, people who go off-script, people who try to talk the agent out of its rules), and pass criteria based on what was written to your systems. Our guide to AI agent evaluation covers how to build that suite.

Before you put one on your phone line

  1. Define the call types the agent owns, and route everything else to a person.
  2. Write the escalation and emergency rules, and make them deterministic.
  3. Script the opening: the agent says it is automated, and that the call is recorded if it is.
  4. Test with accents, noise, interruptions and corrections, and track the 95th percentile response time.
  5. Log every call, review a sample every week, and turn failures into tests.

Frequently asked questions

What is an AI voice agent?

An AI voice agent is software that holds a spoken conversation, usually on a phone line, and acts on it: booking appointments, answering questions or routing calls. It either chains speech-to-text, a language model and text-to-speech, or uses a single speech-to-speech model. The chained design is slower but easier to inspect and control [1].

Yes, with conditions. Outbound automated marketing calls need the recipient's prior consent under PECR regulation 19, and whether a live AI voice counts as "recorded matter" is unsettled [18]. Any call you record or transcribe falls under UK GDPR, so callers must be told how their data is used [24]. Take legal advice for your specific use case.

Do you have to tell callers they are talking to an AI?

For callers in the EU, yes. From 2 August 2026 the AI Act requires people to be informed they are interacting with an AI unless it is obvious, and Commission guidance expects a spoken statement at the start [21][23]. The UK has no equivalent statute, but Ofcom expects it to be very clear [17]. There is no five-second rule.

How fast does an AI voice agent need to respond?

People leave a gap of about 200 milliseconds between turns [8], and a typical chained voice agent replies in about 1.1 seconds [10]. You will not match human speed, so keep replies short, acknowledge before slow lookups, and track the 95th percentile response time rather than the average, because the slowest calls are the ones callers remember.

Can AI voice agents replace call-centre staff?

They can take a large share of bounded, structured calls such as bookings, status checks and out-of-hours enquiries. They are weaker on complex, emotional or high-stakes calls, and benchmarks show reliability falling with accents, noise and self-corrections [5][12]. The realistic goal is shorter queues and better use of your team's time, not a call centre with nobody in it.

What happens when an AI voice agent gets it wrong?

The business is responsible for what its agent says, as a Canadian tribunal found when an airline tried to disown its chatbot [15]. Design for errors: read back critical details, route emergencies by fixed rules, limit what the agent can do, log every call, and hand over to a person, with full context, whenever the agent is unsure.

Talk to us about a voice agent

A voice agent speaks for your business in real time, which is why it needs the strictest guardrails of any agent. We design, build and run voice agents for after-hours and overflow calls, with escalation rules agreed before go-live, callers told they are speaking to an automated assistant, and every call logged and reviewable. See how we build voice agents. If you already run a voice agent and want a named senior engineer to check it, start with a review. Book an audit.

Sources

  1. OpenAI, "Voice agents", OpenAI API developer documentation (accessed 9 October 2026). https://developers.openai.com/api/docs/guides/voice-agents
  2. Lang and Eskenazi, "Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale", 27 February 2025. arXiv:2502.20140. https://arxiv.org/abs/2502.20140
  3. Koenecke et al., "Racial disparities in automated speech recognition", Proceedings of the National Academy of Sciences, 2020. https://doi.org/10.1073/pnas.1915768117
  4. Serditova, Tang and Steffens, "Automatic Speech Recognition Biases in Newcastle English: an Error Analysis", Interspeech 2025. arXiv:2506.16558. https://arxiv.org/abs/2506.16558
  5. Bogavelli et al., "EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents", 13 May 2026. arXiv:2605.13841. https://arxiv.org/abs/2605.13841
  6. Home Office, National Assessment Centre, "Fraud assessment 2025", 9 March 2026. https://www.gov.uk/government/publications/national-assessment-centre-fraud-assessment-2025/national-assessment-centre-fraud-assessment-2025
  7. Information Commissioner's Office, "Five steps to protect your organisation from AI-powered cyber threats", 14 May 2026. https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2026/05/five-steps-to-protect-your-organisation-from-ai-powered-cyber-threats/
  8. Stivers et al., "Universals and cultural variation in turn-taking in conversation", Proceedings of the National Academy of Sciences, 2009. https://doi.org/10.1073/pnas.0903616106
  9. ITU-T, "Recommendation G.114: One-way transmission time", May 2003. https://www.itu.int/rec/T-REC-G.114-200305-I/en
  10. Twilio (Phil Bredeson), "A Guide to Core Latency in AI Voice Agents (Cascaded Edition)", 17 November 2025. https://www.twilio.com/en-us/blog/developers/best-practices/guide-core-latency-ai-voice-agents
  11. Lin et al., "Full-Duplex-Bench", 2025. arXiv:2503.04721. https://arxiv.org/abs/2503.04721
  12. Lin, Chen, Chen and Lee, "Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency", 6 April 2026. arXiv:2604.04847. https://arxiv.org/abs/2604.04847
  13. Twilio (Kahan, Foster, Kovalenko and Maggidwar), "Eleven Best Practices, Tips, and Tricks for using Speech Recognition and Virtual Agent Bots with Voice Calling on the Twilio CPaaS Platform", 13 October 2023. https://www.twilio.com/en-us/blog/tips-speech-recognition-virtual-agent-voice-calling
  14. Koenecke et al., "Careless Whisper: Speech-to-Text Hallucination Harms", ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2024. arXiv:2402.08021. https://arxiv.org/abs/2402.08021
  15. Civil Resolution Tribunal of British Columbia, "Moffatt v. Air Canada", 2024 BCCRT 149, 14 February 2024. https://decisions.civilresolutionbc.ca/crt/crtd/en/525448/1/document.do
  16. OWASP Gen AI Security Project, "LLM06:2025 Excessive Agency", OWASP Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
  17. Ofcom, "Consumers using AI to help manage their telecoms services, but human contact still valued", 25 September 2026. https://www.ofcom.org.uk/phones-and-broadband/telecoms-infrastructure/consumers-using-ai-to-help-manage-their-telecoms-services-but-human-contact-still-valued
  18. The Privacy and Electronic Communications (EC Directive) Regulations 2003 (SI 2003/2426), regulation 19. https://www.legislation.gov.uk/uksi/2003/2426/regulation/19
  19. Information Commissioner's Office, "Telephone marketing", Guide to PECR. https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/guide-to-pecr/electronic-and-telephone-marketing/telephone-marketing/
  20. Information Commissioner's Office, "Warning over robo calls as energy firms fined half a million pounds for unlawful marketing calls", 25 September 2025. https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2025/09/warning-over-robo-calls-as-energy-firms-fined-half-a-million-pounds-for-unlawful-marketing-calls
  21. Regulation (EU) 2024/1689 (the AI Act), Articles 2 and 50, Official Journal of the European Union. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
  22. European Commission, "Quick Facts: Transparency rules for AI systems". https://digital-strategy.ec.europa.eu/en/factpages/quick-facts-transparency-rules-ai-systems
  23. European Commission, "Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of Regulation (EU) 2024/1689", C(2026) 5054 final, 20 July 2026. https://digital-strategy.ec.europa.eu/en/library/guidelines-transparency-obligations-providers-and-deployers-ai-systems
  24. Information Commissioner's Office, "Right to be informed", UK GDPR guidance. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/individual-rights/right-to-be-informed/
  25. Information Commissioner's Office, "Biometric data guidance: Key data protection concepts". https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/lawful-basis/biometric-data-guidance-biometric-recognition/key-data-protection-concepts/
  26. UK GDPR, Articles 22A to 22D, as inserted by section 80 of the Data (Use and Access) Act 2025 (in force 5 February 2026, SI 2026/82). https://www.legislation.gov.uk/eur/2016/679/article/22A

Want it done properly, once? We install OpenClaw isolated, hardened and monitored, then keep it updated under a plain monthly retainer. Fixed setup fee, quoted in writing.

Get set up