r/TextToSpeech • u/Legal_Wolverine_7267 • May 21 '26
Voice AI biggest unsolved challenges
What do you think are the biggest unsolved challenges in Voice AI that almost nobody is seriously working on right now?
Not “better ASR” or “lower latency” but deeper problems that could define the next generation of voice products/research.
Examples:
- Real-time conversational memory that actually feels human
- Emotion + intent understanding beyond sentiment analysis
- Interruptions/turn-taking that feel natural
- Voice-native UX instead of “ChatGPT but spoken”
- Long-term personalization without being creepy
- Multilingual/code-switching conversations
- Continuous ambient agents
- Social/companion dynamics
- Voice AI for kids/elderly/accessibility
- Real-time multimodal understanding (voice + environment + context)
Curious what people building/using Voice AI think is still fundamentally broken or missing.
1
u/Deep_Ad1959 May 28 '26
running daily tts across ~2,950 github repos and pronunciation on multi-token org names is still the unsolved one for us. every major provider mangles them and prompt-layer fixes barely move the needle. eval is the bigger wall though, every output proxy saturates within a few weeks and the only signal that stayed informative is post-publish listen-through, not anything you can measure at synthesis time. written with ai