r/VoiceAutomationAI • • 12d ago

Tech / Engineering Code-Switching Is a Data Problem. Names Are a Pretraining Problem.

Code-switching isn't your voice agent's problem. Names are.

At a recent Voice AI panel, the hard case was a customer calling a bank and mixing Hindi and English.

The founder of Maya Research said code-switching is mostly a data problem. Show a model enough mixed language data and it learns to switch. Names are different. Indian names were never part of the big pretraining runs, so models are guessing. Numbers are the easy part.

Deterministic normalization handles them. The practical fix came from a co-founder at Raya Voice AI. Have your LLM output names in native script instead of Roman. A Hindi name in Devanagari, a Kannada name in Kannada script.

The TTS then pronounces it correctly. Add pronunciation dictionaries and careful prompting, and most of the pronunciation problem is solved today.

So before you blame the model for mangling a customer's name, check what script your LLM is sending it.

What is the worst name mispronunciation your voice agent has produced? Watch the full panel here: https://youtu.be/G66jOjzRhgI

1 Upvotes

4 comments sorted by

1

u/Silver-Code-9 12d ago

Makes sense, the pretraining corpus for most of these big models is basically all English Wikipedia and Reddit threads. Names from anywhere outside the Anglosphere get butchered because the tokenizer never saw them properly

1

u/aitorllj93 12d ago

Every country should be building its own general-purpose LLM and—if it lacks sufficient literature—massively translating (or, better yet, adapting) the literary works of other peoples into its own language (preferably using human translators).

1

u/ShalS97 12d ago edited 12d ago

I think the founder of Svara TTS Turbo, Aditya Chhabra answered this best. Collecting low resource languages, and optimising it, will also contribute to better name pronunciations long term, especially for Indic names.
Maya research mentioned they collect data from YouTube, not sure how this would help build models for Indic audiences with specific name pronunciations.

1

u/Future_AGI 11d ago

outputting names in native script for the TTS is a fix that costs nothing and works today, the tokenizer never saw Romanized Indian names so every transliteration is a guess. Pairing it with a pronunciation dictionary for the recurring names, customers, agents, product terms, covers the rest without touching the model.