Voice AI Hasn't Had Its ChatGPT Moment, PolyAI CTO Says
PolyAI CTO Shawn Wen says voice AI still lacks its "ChatGPT moment" despite billions invested and the arrival of full-duplex speech models.
By Olivia Hart
4 min read
Updated

What's News
- PolyAI CTO Shawn Wen says voice AI hasn't reached its ChatGPT moment despite full-duplex models arriving.
- Investors have poured billions of dollars into voice AI startups, from model makers to meeting note-takers.
- Wen spoke on stage at the HumanX conference last month.
- Otter is developing digital twins that could represent people in meetings.
- Both executives said AI tools should disclose when customers are being recorded or talking to AI.
Investors have poured billions of dollars into voice AI startups, yet PolyAI CTO Shawn Wen says the technology still hasn't reached its "ChatGPT moment."
Wen made the claim on stage at the HumanX conference last month, despite the arrival of full-duplex models — systems that can speak while listening to a user at the same time. His diagnosis: the models can hold a conversation, but they cannot think fast enough to make that conversation feel natural.
"We have reached the milestone of developing full-duplex models. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural," Wen said.
The voice-as-interface thesis has attracted capital across a wide range of categories, from model makers to enterprise customer service providers, and from meeting note-takers to AI-powered dictation tools. New model and tool releases claiming human-like speech arrive weekly. According to Wen, the claims outpace the reality.
What has to change before voice AI breaks through?
For Wen, the gap between demo and deployment comes down to speed and confidence. AI agents in customer service, he argued, should not sound robotic and should give callers enough confidence that they can solve problems.
"I think the next stage will be slightly different because once the voice is good enough, like, and the customer is willing to engage with them for the first two or three turns, they start to build confidence, and over time, they will feel like I probably don't have to talk to a human if the agent can solve my problem," he said.
Alex Gay, CMO of meeting notetaker Otter, pointed to a different bottleneck: understanding who is speaking and what they intend. Speaker identification, intent capture and typing that up with organizational knowledge is, in his view, a key step toward enabling automation.
Otter is also developing digital twins that could represent people in meetings. For that technology to work, Gay said, the output voice must carry the same emotive expressions as talking to a human.
"If you think about the meetings that you're in right now, the best conversations that you have are where you can have debate, and strategic discussions, and when you feel like there's a relationship that underpins it. If you aren't able to have that with an avatar, then it's just a q and a chatbot," Gay said.
Where do today's voice models still fail?
Even as voice AI models have improved, AI assistants often fail to understand users, or meeting notetakers produce the wrong transcript or summary.
Wen said ASR (Automatic Speech Recognition) models often miss important keywords, which creates a problem in capturing the full context of a conversation.
Gay agreed, and framed the stakes for his company in blunt terms. Otter continues to invest in improving transcription, and he identified language as one area where voice models still need to improve.
"For Otter, you know, transcription was never the end point. It was just the layer that we could start to drive some of the productivity gains on the back of. But if your original transcription didn't have the accuracy that you needed, all the follow-up actions that you have become flawed. And the minute that starts to take action, that is wrong. You lose trust in the platform. It is critical for us to continue to improve that ASR model because all of the downstream impacts are significant," he said.
His argument underscores a commercial reality for the sector: transcription accuracy is not a feature but the foundation. When the first layer gets a word wrong, every automated action built on top of it compounds the error — and erodes user trust in the product.
Who tells the customer they're talking to a machine?
The new generation of voice tools also raises transparency questions. Gay said tools should declare to customers that they are being recorded or talking to AI.
Otter wants to build trust among everyone in a meeting, Gay said. Even when the company's bot is not present, Otter wants to try methods such as notifying everyone in the chat that the meeting is being recorded.
Wen said it is equally important to establish that people are talking to an AI on enterprise calls.
The two executives' comments sketch a sector still in its build-out phase: billions invested, full-duplex models shipped, but speed, transcription accuracy and disclosure practices unresolved. Until those gaps close, the breakthrough consumer moment for voice AI — the equivalent of what ChatGPT did for text — remains, by its own builders' account, ahead.
Source: TechCrunch
More from Olivia Hart
Show full bio
Staff writer covering industry trends and analytics at Business Bearings.
593 articles