
← START2 sep · 14 min
START: Aoden Teo, CEO & Co-Founder, Miso Labs: "The most emotive foundation models for voice"
Voice AI can pass a Turing test. For about a minute.
That's a generated clip, though. Have a human actually talk back and the number collapses to six or seven seconds, roughly where generated voice sat three years ago.
One reason, per Aoden Teo of Miso Labs: real conversation isn't turn-based. Around 20% of the time more than one person is speaking, and laughter drives a lot of that overlap, since you laugh at a joke while it's still being told. We also adjust our pacing toward whoever we're talking to without noticing we're doing it.
Voice models struggle with all of this. Full-duplex voice, where a model listens and speaks at the same time, is still extremely early.
So an agent can know your joke is funny and still have to wait until you've finished before it laughs, by which point the timing has killed it.
Aoden describes a second consequence: agents get pushed toward almost "psychotically emotive" behavior. If they can only talk once you've stopped, they need some other way to show they were listening. You finish your sentence, and the thing goes "Hmm?" You've heard it.
Underneath that sits an architecture problem. Voice models have to respond fast, which constrains how large they can be, and fast means something different here than it does in text. Working with an LLM like Claude, Aoden points out, you care how quickly it finishes your code, not how quickly it starts.
Voice inverts that. Nobody needs 10 hours of audio generated in two seconds, because nobody can listen to 10 hours of audio in two seconds; what matters is reaction time. Most architectural decisions trade latency against throughput, and Aoden expects voice to keep moving away from LLM-style designs toward ones built around very low latency.
Miso is already pushing on it. Miso-1 got 3,000 stars on GitHub and 5 million views on Twitter, and they record data in their own LA studio because the internet doesn't contain every kind of audio a voice model might need. Nobody has released a podcast of someone reading millions and millions of email addresses, and people still want voice models that can read email addresses aloud, so teams end up generating some very strange training data themselves.
The clip isn't the hard part. The hard part starts when you talk back.
"The most emotive foundation models for voice"
🎙️Aoden Teo, CEO & Co-Founder, Miso Labs on Fondo START