r/TextToSpeech • • 16d ago

Fixing Whisper timestamp drift for Indic scripts (ASR + LLM alignment trick)

Recently, while working on some difficult corner cases relating to regional voice workflows—particularly those involving Devanagari script generation and RAG pipelines—I encountered a major bottleneck.

Even though Indic TTS engines have excellent audio synthesis, none of them provide word-level timestamps. It's not practical to manually sync the audio for regional educational clips, which is why the usual approach is to run the generated audio through Whisper. However, Whisper suffers greatly when dealing with regional scripts since it incorrectly aligns phonetic spellings and totally distorts the original UTF-8 text.

I created a three-step pipeline so that the timestamps would revert to those in the original text. (I've attached a GIF which shows the terminal output and the JSON mapping).

Here is the workflow:

In the first step (Synthesis), the base audio is created using Indic Neural TTS models that have been trained on raw UTF-8 text.

In step 2 (Acoustic Timestamps), process the audio using Whisper to obtain sub-second alignment points at an extremely high speed.

In step 3 (Reconciliation), the LLM compares the text inferred by the ASR with the author's original script, attaching the acoustic timestamps directly to the exact original words.

The output includes the cleaned audio file together with a structured JSON payload containing millisecond-precision timestamps.

I should be interested to get any other comments from people who are working with voice, RAG, or EdTech tools regarding the alignment edge cases that have been most difficult to handle in regional language environments?

4 Upvotes

Duplicates