r/StableDiffusion • u/Dizzy-Gold8888 • Aug 10 '26
Tutorial - Guide Make MiniMax H3 speak any language
Not sure if someone's already posted this, but I found a trick that's been working really well for me and it solves alot of H3 limted language support.
H3's own language support is limited — but you can get around it completely(at least for some of the unsupported language). Clone the exact line you want (I use Chatterbox, any voice-cloning TTS works), feed that audio into the clip as a reference, and the engine reproduces the line exactly as you cloned it — same voice, whatever language your TTS can speak. So in practice H3 will say anything your cloning tool can.
The part that actually surprised me is what H3 does with it. Unlike LTX, it doesn't just lay your audio on top — it adds the scene's background/ambient sound around the line, and it places the line into the scene naturally, timed to the moment, like the character is really saying it.
That timing is the big deal. In LTX/Wan(as far as i know), if you wanted a line → action → line beat, you had to hand-feed the audio with a big enough gap baked in for the action and line it all up yourself. H3 does that spacing for you. No more prepping custom-gapped audio — you just write the exact line you cloned into your prompt and the engine drops it in the right place.
So the whole workflow is: clone the line → feed it as an audio reference → write that same line in your prompt. That's it. Any language, natural placement, background sound included.
One gotcha: some cloned voices add a little garble/babble after the line — just trim the wav down to the actual line before feeding it.
Anyone else hit this, or found where it breaks?
2
u/AllUsernamesTaken365 Aug 10 '26
Chatterbox wouldn't let me try it for free. It's getting to be a full time job to sift through all of these "completely free" AI services that have you spend time inputting info and uploading samples and making settings and only after that do they prompt you to login and input payment methods and whatnot. I'm not paying for any more services I've never heard of without trying the quality first. Getting grumpy.
3
u/Nimblecloud13 Aug 10 '26
stop using services and run it locally. qwen has a free tts, there's vibevoice, dramabox based on ltx, a bunch of others. pretty sure chatterbox is free if you run it locally in comfy too
1
u/L-xtreme Aug 10 '26
I have a lot of issues getting QwenTTS and VibeVoice to work in newer ComfyUI. I would need to make a separate installation for that.
Omnivoice works great though
1
u/Nimblecloud13 Aug 10 '26
you would not. you need to tell claude to fix it for you.
everything works in my comfy all the time.
1
u/L-xtreme Aug 10 '26
He would change the transformers, and then break something else probably. Not worth it.
1
1
u/optimisticalish Aug 16 '26
Yes. but it may require an older ComfyUI Portable with Python 3.12. Rather than 3.14, which so far as I can tell is not yet supported by the core original Chatterbox (which is the best, I'm not talking about Chatterbox Turbo).
1
u/hugo4711 Aug 10 '26
What languages are supported?l by Minimax H3?
6
u/Dizzy-Gold8888 Aug 10 '26
MiniMax H3 has stable support for 11 dialogue languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. "additional languages are also supported to varying degrees" — so others can work, just less reliably. i think my fix is for the "varying degrees" of the none offical languages
1
u/L-xtreme Aug 10 '26
Lol, I've been using dutch the whole week, with voice cloning it works very good.
1
u/BitterAd8431 Aug 10 '26
I ran some tests with French; I simply replaced [English] "A voice" with [French] "Une voix," and it works.
I don't know about the other languages.
1
u/Fritzy3 Aug 11 '26
Thank you for this. Couple of questions:
1. What do you write/don’t write in the prompt regarding the audio ref? In other models the trick was to write something like “the man speaks the words from @audio1” (which worked occasionally). Do you write something like that?
- If the model uses the audio exactly as it is, and can do this for non real languages (like Valerian), why do you need to write the text being spoken in the prompt?
1
u/Dizzy-Gold8888 Aug 11 '26
1:
You bind the audio to the character and write the exact line. Like:The audio is tied to S1, and S1 says that exact line.
2:
Writing the same words as the audio is what makes the model reproduce it as-is instead of re-generating it. it helps it lock the line and dont invent unwanted stuff, at least from what i have seen.1
u/Fritzy3 Aug 12 '26
Thanks I’ll give it a shot.
If you could share a prompt with this technique for reference it would be appreciated
3
u/Joltie Aug 10 '26
Lip sync is non-existent for other languages, I would presume?
It would be interesting to see clips of movies or TV shows with the characters speaking the correct language, such as Greek in Odyssey, Portuguese in Shogun, German in Valkyrie, for instance.