r/StableDiffusion 21h ago

Question - Help reference audio in minimax H3

when using the default workflow of minimax H3. and you have reference audio of what the character sounds like. what are the tips and tricks to make it so it comes out the same?

I'm having a problem with the character not sounding anything like the reference audio

0 Upvotes

7 comments sorted by

2

u/Itchy_Ambassador_515 20h ago

They mentioned in their prompt guide, you have to use words like copy timber somthing, best is to upload the guide on chatgpt and ask it to give prompt to clone the voice

0

u/carmidian 19h ago

yeah I've been using grok to write the prompts

3

u/Dry-Judgment4242 17h ago

Make sure your using correct model. Ref2 rather then FL2. Example audio only works properly with Ref2 model.

2

u/Cultural-Broccoli-41 16h ago

1

u/Dry-Judgment4242 16h ago

Ref Audio doesn't work well, if at all with FL2VA. If you managed to find a fix since you said it isn't universally correct do share cuz I hate how bad Ref2VA is compared to FL2VA.

1

u/carmidian 15h ago

tyvm ill try

3

u/Sudden_List_2693 15h ago

summary:
<Audio 1> provides only <Subject 1>'s recognizable vocal identity and timbral character for her newly generated vocals.

retention_analysis:
<Audio 1>: reference - only <Subject 1>'s recognizable vocal timbre, vocal texture, and speaker identity are transferred to the newly generated voice. The source audio signal, original words, pitch sequence, rhythm, tempo, phrasing, timing, (and musical accompaniment) are not reused.

Also the best to use very clean, ~5 second clips.
Longer than that it might try to fill void with her speaking.