Neil Weldon
Summary: Agentic AI fails in the part of the stack that’s hardest to see. The agent can misread the caller or invent an answer. The voice estate underneath it can drop, misroute or degrade the call first. Both look the same from the outside in, so test the agent and the call together.
Once an agent is live, the failures are the hard kind to chase. When calls start to drop, for example, someone may report it a day later, and by then the evidence is gone. A voice agent can fail in more ways than a chatbot or a phone menu, and you may never know why. That’s what my upcoming Genesys Xperience talk is built on.
Why do customers pick up the phone?
ContactBabel’s 2026 Decision Makers’ Guide asked a thousand consumers which channel they’d choose. Calling won whenever the problem was urgent, emotional or complicated, across every age group and income bracket, and has done so for a decade. People call when it matters. Yet Invoca’s study of more than 60 million phone calls found that only 56% of inbound US calls involve speaking with a human. The other 44% hang up, hit voicemail, land out of hours or get lost. Those are the calls that mattered most to the person making them, and the gap that agentic voice is meant to close. It works when it works: Service 1st Federal Credit Union added a voice AI layer and cut its call abandonment rate by 96%.
The Four Horsemen of AI voice agent failures
Assembly AI surveyed more than 450 voice agent builders this year. Around 82% of them felt confident in building agents, while about 75% hit reliability barriers in production. That gap is the whole problem. In my Genesys Xperience session, I break it into Four Horsemen of AI Voice Agent Failures, the failures that keep turning up together.
-
Audio quality: Someone calls from a busy street, the line is poor, there’s a delay, or they talk over the agent.
-
Service quality: The agent misreads what the caller wants, invites an answer, or treats someone who is clearly upset like someone who isn’t.
-
Compliance and security: Customer data has to stay protected, the agent has to sound like your brand, and a caller shouldn’t be able to talk it out of its own rules, as we’ve seen in very public cases.
-
Complexity: Every engine sounds good in a lab, but drive twenty minutes down the road in Europe and the accent changes, and so does what the agent hears.
And the fifth, the one nobody names
Underneath all four sit the problems we’ve always had with phone systems. Manual testing, a few sporadic calls at a time. Finding out something broke because a customer told you. No way of knowing whether your numbers even ring in other countries. No evidence of what went wrong, so everyone blames someone else.
These are the ones I keep coming back to, because they reach the agent looking exactly like agent faults. A caller in one market gets audio so degraded that recognition gives up. The transcript shows an agent that understood nothing, so the model gets rebuilt while the carrier problem that caused it stays where it is.
Telling them apart is more ordinary than it sounds. Place a test call from inside the country the customer is calling from, and you will know whether the number rang, if the audio arrived clean, and whether the transfer to an agent went through cleanly. If all three pass and the caller didn’t get what they needed, it’s the agent. If they don’t, no amount of prompt work saves you. You can’t safely deploy AI on a voice estate you have no visibility into, a point my colleague Christine Ramsey made well in Smart Customer Service.
Who owns AI voice agent behavior?
Usually nobody, which is the uncomfortable part. Gravitee found that 85% of businesses have nobody formally accountable for how their AI agents behave, while 81% of teams feel pressure to launch fast. Moving fast without ownership is how a small problem can quickly become a pretty serious one, and customers blame you regardless. Invoca found consumers are 3 times more likely to blame the brand than the AI.
What does the session cover?
Mostly, it will cover how to test in a way that survives real callers. The early approach was a few manual calls by hand before launch, then waiting for complaints to tell you what was broken. That doesn’t hold water once an agent is taking thousands of calls in a dozen countries, and especially if you have “too big to fail” numbers that, if down, are financially detrimental to your business.
I’ll walk through the five steps we see working. Governance, as above. Connection, so the call arrives and passing it to a person doesn’t lose the caller or what they’ve already said. Quality, so the agent resolves the query. Compliance, so customer data stays safe and the agent stays on script. Complexity, so languages, accents, noise and volume get tested the way callers arrive.
The same problem sits behind all five. Test once and you’ve proved it worked once. Testing before launch isn’t enough when your models, your carriers and the rules keep changing, so the testing has to be continuous too.
Common questions
How do you evaluate a voice agent?
You evaluate the agent and the call it arrives on. Check whether the agent understood the caller and resolved the request, and whether the call arrived clean, connected in every market and survived the handover to a person. Score only the transcript and you miss half the failures.
How do you know if a voice agent is ready?
When you have evidence rather than a good demo. Test calls in-country show whether the number rang, whether audio arrived clean, whether the agent understood and whether the handover worked.
Session Details
Before the First "Hello": What Makes Agentic Voice Succeed
Thursday 3rd September, 1:10 PM, Showcase Theater 2, Wynn Las Vegas. Register here.
Klearcom is an Official Silver Sponsor at Genesys Xperience 2026. Add the session in the Xperience app, and scan the QR code at the end of the talk for the one page summary and early access to our AI agent testing.
