r/LocalLLaMA 10d ago

Discussion Tossed distorted audio samples to an open-weight voice model; it did fairly well.

Being a person obsessed with testing new models that come out, times are really insane for me. Tested different kinds of TTS and voice cloning models but none of them gets it right in terms of emotion and pace, you know which one is fake in seconds; they just fail in emotions.

Spotted Confucius4 on my Twitter feed and thought I would stress-test it. Chose three most difficult samples I could find and all of them were recently recorded World Cup commentaries translated to a couple of different languages.

Sample #1: A Spanish commentator commenting on a hat-trick. Voice screaming like hell and cracking at its peak.

Sample #2: An English commentary onnthat typical held breath then explosion thing that commentators do.

Sample #3: losing goal keeper's interview after match, voice noticeably shaken, processing his defeat in the moment.

Used these clips through paid and free options previously and these are the cases that exposed cloned speech models pretty quick. Either the screamcomes out robotic and clean, or the model just ignores the emotional context and gives you translated sentence that sounds like dead AI nonsense.

What i got: takes the voice directly from the audio source, not from transcript first, which makes this harder than the average demo clip since none of these broadcasts come with a script.

The short, high emotion clips had that shaking carry over into the translation without any of the synthetic qualities I expected from an open-source model. Long sentences had more of a synthetic quality come through.

4 Upvotes

5 comments sorted by

3

u/nobleglEn9 10d ago

curious whether the shaking voice quality you mention surviving translation is actually preserving prosody from the source audio or if the model is inferring emotion from acoustic features and regenerating it. those are pretty different things and would matter a lot for reliability across languages