Talk to your phone, and it talks back. Ask a customer-service line for help, and a machine may now be the first voice you hear. Dictate a message, and an AI turns it into text before you have finished the sentence. Voice AI is, by any measure, everywhere — and billions of dollars in venture funding are betting that speaking will become the next dominant way humans interact with computers.
Yet according to some of the people building that future, the technology has not actually arrived yet. Speaking at the HumanX conference, executives from two of the most prominent companies in the field — PolyAI's chief technology officer Shawn Wen and Otter's chief marketing officer Alex Gay — argued that for all the weekly product launches and human-sounding demos, voice AI is still waiting for its breakthrough moment: the equivalent of what ChatGPT did for text-based AI in late 2022.
The milestone reached — and the one that matters more
Wen's argument begins with what has already been accomplished. The industry has, he says, reached the milestone of full-duplex models: systems that can speak while simultaneously listening to the caller, the way humans naturally interrupt, hesitate and overlap in real conversation. That is a genuine technical leap from the clunky half-duplex systems of the past, where a caller had to wait for a prompt before answering, and it explains why the current generation of voice assistants sounds so much more natural than its predecessors.
But full-duplex speech is only the beginning, Wen argues. The next and more important challenge is making reasoning very fast — fast enough that the model can fetch answers and assemble responses without the awkward pauses that still betray its artificiality. When a human customer-service agent is asked a question, the answer arrives in a continuous flow; with most AI voice agents, there is still a lag, a silence, a tell-tale beat while the system thinks. Until that latency disappears, Wen suggests, the conversation will never quite feel natural.
This is a subtler problem than it sounds. Text-based chatbots can hide their thinking time behind a typing indicator or a partial response. A voice agent has nowhere to hide: silence on a phone call is conspicuous, and a voice that hesitates at the wrong moment sounds confused even when it is merely computing. The bar for a voice AI to pass as a competent conversationalist is arguably higher than for a text one — it must think at the speed of speech.
Sounding human is not enough
There is another trap the industry is still falling into, both executives suggested: equating a good voice with a good conversationalist. Every week brings a new model or tool release claiming to sound human, and to be fair, the synthetic voices have improved enormously. But a human-sounding voice that misunderstands you is arguably worse than a robotic one — it raises expectations it cannot meet.
Wen's view of customer-service AI reflects this. His company's agents should not sound robotic, but sounding natural is in service of a larger goal: giving callers enough confidence, in the first two or three turns of a conversation, that the agent can actually solve their problem. Confidence, not vocal fidelity, is the currency that matters. Once callers trust that the system can handle their request, Wen argues, they will stop feeling the need to be transferred to a human — and that is when voice AI becomes economically transformative for enterprises.
Alex Gay, speaking for meeting notetaker Otter, made a parallel point about a different setting. The company's ambitions go beyond transcription into automation: speaker identification, intent capture, and layering that understanding onto an organisation's knowledge so meetings can drive real work. The company is also developing digital twins — AI representatives that might stand in for people in meetings they cannot attend. For that to work, Gay says, the output voice has to carry the same emotional expression as talking to a human in a meeting.
His reasoning cuts to the heart of why voice AI is so hard. The best conversations people have at work, he noted, are debates and strategic discussions underpinned by a relationship — a sense of trust, rapport and shared context. If an avatar or agent cannot participate in that kind of exchange, it remains, in his words, just a question-and-answer chatbot with a pleasant voice. The voice is the surface; the relationship is the product.
The unglamorous bottleneck: transcription
Beneath all the talk of duplex models and digital twins lies a deeply unglamorous problem that both executives kept returning to: the AI assistants often simply do not understand what is being said. Automatic speech recognition — the ASR layer that converts raw audio into text — still misses important keywords, and when the transcript is wrong, everything built on top of it goes wrong too.
Wen flags keyword-missing as a core failure: a model that drops a crucial term cannot capture the whole context of what a caller wants. Gay is even blunter about the stakes. For Otter, transcription was never the end point — it was just the layer on which productivity gains could be built. But if the original transcription lacks the accuracy the product needs, all the follow-up actions become flawed. And the moment an AI takes an action that is wrong, the user loses trust in the platform. The downstream impacts, in his phrase, are significant.



