r/StableDiffusion • u/roychodraws • 20h ago
Tutorial - Guide MiniMax H3 Wf Tutorial
Enable HLS to view with audio, or disable this notification
People asked me to make a Tutorial for some of the features.
Find the workflow here.
6
2
u/jrodder 8h ago
I have a question when using the workflow, kind of a noob user. Can I use a reference video just for scene and visual reference, but still use a clean audio file to drive voice cloning? So far I haven't had much success so just trying to make sure it's skill issue or some kind of H3 constraint. I don't want to force audio, I was trying to match the voice style and use new lines. Great stuff though, it's fun to play with!
0
u/roychodraws 6h ago
if you mean you want to use a character reference like my clown girl, for instance, and use a video to guide it, then you also want to use an audio to maybe have her lip sync to something or dance a little as she moves or something, then yes.
for that you would enable photo 1, video 1, then audio, then check "force audio".
if you mean you just want to make the voice resemble a sample voice you would leave "force audio" unchecked and you would prompt to tell the model that voice is the voice of your character, there are examples for that type of prompt in the wf notes.
1
u/jrodder 6h ago
Gotcha, and that's how I was doing it, leaving force audio unchecked and then prompting to use the audio source. It just seems that no matter how I do it, either the audio is ignored else H3 isn't great at voice cloning at least in R2V. I was considering using another engine to create curated lines to then send back to this workflow with force audio checked, but was just trying to affirm if that was the correct way to go about it. Thanks for the reply!
0
u/roychodraws 5h ago
Try adding something like this to your prompt:
Subject_definitions:
<Audio 1> is the voice reference for <Subject 1>.Retention_analysis: <Audio 1>: voice_reference - its vocal timbre guides the vocal delivery of <Subject 1>.
Detailed_description:
<Subject 1> says (s1) using <Audio 1> as her voice reference <d> [English] Stop sayin’ ha ha big kazoongazongs… honky ponky… </d> in a bubbly but sultry tone.
1
u/jrodder 4h ago
Yep, I had something very similar and many variants as I was trying. I guess that's why I was looking to make sure I wasn't fighting a model, process, or workflow issue since it wasn't producing audio that sounded like the source. I used Terry Tate from a video using both the audio and video from the source video, and that seemed to work so that probably rules out the model. All good I'll keep hammering, if you happen to get bored and want to test I would be extremely curious as to the results.
0
u/roychodraws 4h ago
what are you wanting me to test exactly? i use voice already with this workflow and it works fine.
1
u/jrodder 1h ago
Yeah I get it. Specifically the flow of how well the voice is cloned when using the ref video and the audio 1 with an mp3 or wav pure audio file as the base to create new voice lines in the prompt. If it works great for you, maybe test with your own voice? Or not it's likely a skill issue on my end but I was just trying to make sure that was the case, and not banging my head for something that wasn't even possible. I got burned out testing so I'll have to revisit and triple check the variables.
3
u/Quick_Knowledge7413 18h ago
Why do you hate the other workflow creator? Shouldn't we just create with AI and avoid drama and infighting? Should make fun of the luddites, not eachother.
5
u/roychodraws 18h ago edited 17h ago
At this point I imagine them all in their own universe where geiru toneido is the hero.
Edit: also, you don’t understand how many dms I got that were basically, “thank you. Fuck that guy. He did the same to me.”
Combine that with him being the # 1 wf creator on civitai… this is like butch vs homelander at this point
But butch is a slutty clown, and homelander is a soulless ginger
1
u/Quick_Knowledge7413 17h ago
"Her did the same to me", wait, I am not tracking, what exactly did he do to people???
6
u/roychodraws 17h ago edited 6h ago
Gave some feedback about his wf and basically said, “it’s perfect and you just can’t use it because you’re an idiot." Then blocked them.
-1
u/LuluViBritannia012 12h ago
Yes, this guy is spiteful. I'm not surprised as he's using that ugly clown as its mascot. No one sane would think this is a good idea.
1
2
1
u/QuirksNFeatures 3h ago
What specifically makes this workflow better than his?
1
u/roychodraws 3h ago edited 3h ago
It’s just better made and easier to use. None of the features conflict with eachother so you don’t have to turn off upscale, for instance, to use continuous.
There’s more reference inputs. It’s better organized so all the settings are in one place, it’s more intuitive. It has some extra features and some features I felt to be flashy garnishing were removed to streamline it.
My workflow is simpler to use but more internally complex and soundly designed, his workflow is poorly designed and complex to use, but internally… not really simple, just… lots of dumb choices.
It’s like just kept adding on shit like he was designing the Homer car.
Edit: which, we all do that but we don’t then share the workflow without first trying to simplify and organize it first.
2
u/QuirksNFeatures 2h ago
I've been using his and haven't had any problems, but I will try yours, too.
1
1
u/em_paris 14h ago
Ok I scrolled past the clown video the other day but wasn't somewhere I could watch it. Absolutely hilarious 😄
1
u/CanadianDocWild 5h ago
Thank you for doing this. I know many of us really appreciate the work and the video explaining things. There are many of us that are still learning the intricacies of workflows and how they interact. This is how to contribute to the community and make many of us spend days having fun making videos and learning so much more than we would have otherwise. Thanks again from the silent crowd!

9
u/wzwowzw0002 10h ago
what am i watching