What breaks when AI voice agents go live?

"What breaks when AI voice agents go live, and why almost nobody sees it coming?" Ilya Ostrovskiy, Chief Product Officer at Tenios and Apifonica, took to the stage at CSVD 2026 in Hamburg on September 8th.

Darija Fjodorova

September 21, 2026

7 min read

“What breaks when AI voice agents go live, and why almost nobody sees it coming?” Ilya Ostrovskiy, Chief Product Officer at Tenios and Apifonica, took to the stage at CSVD 2026 in Hamburg on September 8th.

The AI industry spends a lot of time talking about how to launch an agent – which model to use, how to build the conversation and how to get a pilot running. But production starts when the pilot ends. Here are five failures that tend to appear only once real customers start calling. These failures are rarely caused by the language model itself. They happen in the systems, processes and infrastructure around it.

The agent says it did something. The system says otherwise.

-> A caller asks an agent to update their address.

<- The agent responds that the address has been updated.

The transcript looks perfect, but the call never happened. The CRM record remains unchanged. This is one of the most dangerous production failures because conventional quality checks can miss it. The conversation sounds correct even though the actual business process failed.

The same problem can appear when an agent reads the wrong field from an external system. One example involves a BMW dealer in Toronto, where an AI agent interpreted a loan balance as a purchase price. The language was correct. The data interpretation was not.

The lesson is straightforward: a successful interaction should not be measured by whether the agent said the right thing. It should be measured by whether the intended transaction actually happened.

Authentication works in the test environment. Then real callers arrive.

Voice AI has a difficult problem that is easy to underestimate – authentication. A customer number, IBAN or other identifier can contain many characters. One incorrect character can invalidate the entire transaction.

Recognition accuracy can fall sharply as inputs become more complex, moving from pure digits to mixed alphanumeric identifiers and full customer numbers.

A studio microphone is not a mobile phone in a car. A clean test recording is not a caller speaking through an 8 kHz telecom codec with background noise, frame loss or muffling. What sounds manageable in a room becomes much harder for a voice system operating over a real telephone connection.

The implication for testing is important.

Do not test the agent only through a laptop microphone in a controlled environment. Test it on the actual phone path, with real recordings and the types of callers it will serve.

The handover works. The customer context does not.

Getting a customer to a human agent is often treated as a Voice AI feature. In reality, it is a telephony problem too.

When an AI agent hands a call to a human, what happens to the context? Does the human agent know why the customer called? Does the customer have to explain everything again? Does the transfer work consistently when the AI sits on top of another telephony environment?

One of the areas where the infrastructure beneath the AI becomes critical is handover. Research shows that many organisations still fail to carry context across channels, while customers are highly frustrated by having to repeat themselves.

Escalation to a human is a telecom capability, not a prompt. That distinction matters because changing a prompt cannot fix a handover architecture that does not carry the required context.

Nobody notices the failure.

Perhaps the most uncomfortable examples are the cases where an AI system had already been failing for months or even years. The issue was not that the failure existed, it was that the organisation running the system did not find it.

One case involved a government service where callers selecting a Spanish-language option received an English script. The issue remained live for months and was eventually exposed by a customer on TikTok. The attempted fix then created another problem for English callers.

If customers are the first people to discover that your AI is broken, your monitoring strategy has failed. Zero rollbacks are not necessarily a positive signal. Organisations with more mature guardrails can actually report more rollbacks because they detect problems earlier.

A system that has never been rolled back may simply be a system nobody is watching closely enough.

The model is not your bot.

When companies evaluate Voice AI vendors, the discussion often centres on the model: Which model is being used? How fast and accurate is it? How does it compare with the latest model?

But the model is only one part of the production system. The caller experiences the entire chain – the phone line, carrier, network, integrations, tools, APIs, handover logic and the model itself.

Latency makes this particularly visible. Available response-time budget disappears once the model is combined with the rest of the production stack. Realistic tests across realtime systems showed response times reaching roughly 900-1,150 milliseconds. At that point, the model may be performing well while the overall interaction still feels slow to the caller.

The same applies to resilience.

If an upstream model provider goes down, what happens to the phone call? If the answer is that the call stops, then the production system has a dependency that sits outside the organisation’s control. The Voice AI stack is a supply chain rather than simply a bot.

The five failures have one thing in common

The five cases come back to one conclusion – none of the issues is primarily a model problem.

A transaction can fail even when the language is perfect. Authentication can fail because of the phone environment. A handover can fail because the telephony layer cannot carry context. A serious issue can remain invisible because nobody owns the monitoring. A perfectly capable model can become unreliable when the surrounding supply chain breaks.

Three of these five failures cannot be fixed from above the phone line. That changes how Voice AI projects should be evaluated. The question is no longer only whether the model can understand and respond, it is whether the entire production system can execute, authenticate, transfer, observe and recover.

What was changed after the failed pilots

First, we start with real call recordings rather than a polished demo. Analyse which cases occur, how frequently they occur and where an AI Voice Agent can genuinely take over.

Second, test the transaction rather than the sentence. Acceptance should mean that the expected state actually changes in the CRM, billing system or ticketing environment.

Third, telecom infrastructure is a part of the Voice AI product. We are a telecom operator, rather than an AI company that simply rents a phone line. Our production approach includes control of the underlying telecom infrastructure and AI-layer failover. These are architectural decisions, not prompt improvements.

Five checks to make on Monday

  1. Pick one flow and check whether the record in your system actually changed.
  2. Call your own agent from a mobile phone and test authentication with a real customer identifier.
  3. Ask for a human and measure how long the handover takes and how much information you have to repeat.
  4. Find the person who actually owns the AI process and knows which production metrics are being monitored.
  5. Ask the vendor what happens during a major model-provider outage and ask to see the runbook.

A Voice AI system does not operate in a clean demo environment. It operates inside a telecom network, connected to business systems, used by real people and expected to deliver real outcomes.

That is where the failures become visible. And that is where production readiness actually matters.

Share

Blog

Latest insights from Apifonica

View all posts

Talk to our experts
Schedule a meeting to discuss your communication setup and find the right starting point
After submitting the form, you will receive a response within one business day.
30 minute discovery session
A focused discussion on your current communication setup, priorities and where automation or infrastructure could help.
Product overview
A walkthrough of the platform areas relevant to you - AI Agents, Messaging, Voice or Numbers.
Next steps
A clear recommendation on scope, timeline and what a pilot or migration would involve.
Contact form person