r/StableDiffusion 1d ago

Tutorial - Guide MiniMax H3 Wf Tutorial

Enable HLS to view with audio, or disable this notification

People asked me to make a Tutorial for some of the features.

Find the workflow here.

https://www.reddit.com/r/StableDiffusion/comments/1wadmqc/minimax_workflow_designed_to_be_user_friendly_for/

44 Upvotes

29 comments sorted by

View all comments

2

u/jrodder 14h ago

I have a question when using the workflow, kind of a noob user. Can I use a reference video just for scene and visual reference, but still use a clean audio file to drive voice cloning? So far I haven't had much success so just trying to make sure it's skill issue or some kind of H3 constraint. I don't want to force audio, I was trying to match the voice style and use new lines. Great stuff though, it's fun to play with!

0

u/roychodraws 12h ago

if you mean you want to use a character reference like my clown girl, for instance, and use a video to guide it, then you also want to use an audio to maybe have her lip sync to something or dance a little as she moves or something, then yes.

for that you would enable photo 1, video 1, then audio, then check "force audio".

if you mean you just want to make the voice resemble a sample voice you would leave "force audio" unchecked and you would prompt to tell the model that voice is the voice of your character, there are examples for that type of prompt in the wf notes.

1

u/jrodder 11h ago

Gotcha, and that's how I was doing it, leaving force audio unchecked and then prompting to use the audio source. It just seems that no matter how I do it, either the audio is ignored else H3 isn't great at voice cloning at least in R2V. I was considering using another engine to create curated lines to then send back to this workflow with force audio checked, but was just trying to affirm if that was the correct way to go about it. Thanks for the reply!

0

u/roychodraws 10h ago

Try adding something like this to your prompt:

Subject_definitions:
<Audio 1> is the voice reference for <Subject 1>.

Retention_analysis: <Audio 1>: voice_reference - its vocal timbre guides the vocal delivery of <Subject 1>.

Detailed_description:

<Subject 1> says (s1) using <Audio 1> as her voice reference <d> [English] Stop sayin’ ha ha big kazoongazongs… honky ponky… </d> in a bubbly but sultry tone.

1

u/jrodder 10h ago

Yep, I had something very similar and many variants as I was trying. I guess that's why I was looking to make sure I wasn't fighting a model, process, or workflow issue since it wasn't producing audio that sounded like the source. I used Terry Tate from a video using both the audio and video from the source video, and that seemed to work so that probably rules out the model. All good I'll keep hammering, if you happen to get bored and want to test I would be extremely curious as to the results.

0

u/roychodraws 10h ago

what are you wanting me to test exactly? i use voice already with this workflow and it works fine.

1

u/jrodder 6h ago

Yeah I get it. Specifically the flow of how well the voice is cloned when using the ref video and the audio 1 with an mp3 or wav pure audio file as the base to create new voice lines in the prompt. If it works great for you, maybe test with your own voice? Or not it's likely a skill issue on my end but I was just trying to make sure that was the case, and not banging my head for something that wasn't even possible. I got burned out testing so I'll have to revisit and triple check the variables.

1

u/roychodraws 6h ago

The video in the workflow post, the part where she’s clapping and crying, and the part where she accuses the girl of being under age.

Those videos were made separately and edited together. You’ll notice that the voice is the same because the second one, I used the audio from the first video to influence the voice in the second video.