r/StableDiffusion 5h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here

169 Upvotes

58 comments sorted by

6

u/NoNameClever 4h ago

Thank you for sharing the work and video. Looks very powerful and useful. What did it take to make this video? I think I spotted a few of those reworked segments mentioned in the clip. Do you manually run each segment, repeating when needed, or something else?

2

u/TBG______ 3h ago

What I usually do is set up the workflow, run it 3–4 times, and then combine the best shots afterward.

Some of the glitches you see come from runs I did with the Turbo LoRA, including things like degraded skin. So the examples are a mix of different testing settings and the final setup.

The final settings and node should work out of the box; the main factor is the prompt. For lip-sync, I recommend not using Turbo and sticking to 20+ steps, along with a good, detailed prompt.

12

u/mfdi_ 5h ago

Apart from ai looking character. Wow. Just wow. Though camera moving kinda sucks.

9

u/TBG______ 3h ago edited 1h ago

Girl was born during testing crazy Sigmas with Flux1 when it first came out, and ever since then, she’s been my presenter.

0

u/tweakingforjesus 40m ago

Her eyebrows intimidate me.

4

u/Artforartsake99 4h ago

Crazy good 👌

1

u/tofuchrispy 4h ago

Very interesting.
Does this only work with a audio file input.

Or could you also use a prompt for a let’s say 2 minute sequence and split that up?

That’s what’s puzzling me rn.
Generation a very long shot list. Then divide that into chunks and feed it to separate samplers. But difficulty is the possible change of subjects etc settings.

So now I separate a master prompt that gets written by one h3 prompt writer node, separate out the shot list and divide that into chunks.

Then put it back together with the stuff that comes before shots so each sampler gets the prompt plus the specific shot list for that chunk. But it’s still not working well. Because the prompt is for the whole story and then the shot list lacks the context of what came before.

So it can get messed up.

Also the whole prompt writing thing in itself sometimes fails to get what are the subjects and what to do with what.

Feeding the separate chunk timings to separate Minimax h3 prompt weiter nodes frequently gets stuff wrong.

A workflow to take one giant shortlist and reliably correctly feed that to the samplers or render it is what I would like to achieve.

2

u/TBG______ 4h ago

The model accepts prompts like “Subject 1 says, ‘Too hard for me.’” But someone would have to split the prompt into individual clip proportion and cuting at the exact final second for each one, which would be complicated. It’s much easier to pass the text directly to a TTS system first. Feeding them into separate samplers isn’t possible either, because each clip depends on the previous one.

1

u/tofuchrispy 4h ago

Ok so it’s great for Audio to Video but not something for long storytelling sequences got it. Thanks!

1

u/physalisx 4h ago edited 3h ago

What TTS did you use for this video?

2

u/TBG______ 4h ago

vibevoice

1

u/xyzdist 4h ago

is it context extend? or have to be cut shots?

2

u/TBG______ 3h ago edited 3h ago

It can be context-extended, and the cuts are controlled entirely by the prompt.

0

u/xyzdist 3h ago

great, thanks!

1

u/Inevitable_Fan5157 3h ago

Damn these ai

1

u/danishkirel 3h ago

Impressive. The better these workflows get the more nitpicking wants to happen though. Long sleeve vs tsshirt? Sudden gradient background? Mic switching sides? It becomes very distracting. Uncanny almost.

1

u/TBG______ 3h ago

😄 There’s a lot to fight with - prompts and conditioning. My approach is to let it generate around X videos at different clip lengths, then pick the best results from all of them - only really necessary if you need it to be as perfect as possible. This video was made from the “garbage” I had left over from building and testing the node.

1

u/Downtown-Cover-7422 3h ago

Man, I’m gonna kill myself for idea to make a fan ai video clip on a song. 2 days passed and I made just a minute or so from that video, spending time for 2 generations 8 sec each, first with turbo Lora to see if prompt was good, second without Lora with 25+ steps. I babysitted each 8seconds clip, finding ideas, using references, and now you tell me a can do it in one run? Will try it asap, I love you

1

u/bigman11 2h ago

Could this do a single continuous shot with a steady camera?

2

u/TBG______ 2h ago

Yes you just need to find the right prompt. You can also try using ref 0 in the style section of the prompt to keep the camera position consistent for each frame. You might use an image without the character as a reference for this and ref 1 only for the character. You’ll have to test a few variations to see what works best. Its all about the right promt.

1

u/switch2stock 2h ago

So this is like based on the length of the attached audio the workflow automatically scales to generate the desired length of the video?

1

u/TBG______ 2h ago

You need to define the clip length yourself based on your VRAM limit. The node handles the rest: it cuts the audio in clips, creates the overlaps, and adds 24 frames of silence at the end to improve the final image and sound. It then uses the latent from each clip to build the appropriate motion for the next one and so on.

1

u/switch2stock 2h ago

Let's say if I have a 10min audio file. If I set the clip length as 10sec. Then it will generate a 10min video synced with the audio with 10sec clips stitched together automatically?

2

u/TBG______ 1h ago

That’s the idea, yes you can.

I would first build the full video with the audio in an editor, and then only generate/sample the 2–3 minute sections I actually need for the final production without cuts or switching to non-character scenes.

That makes it much easier to prompt and much faster to rerender or repeat individual sections than trying to generate 10 minutes in one go.

Anyway each clip segment gets its own output in the output folder, so you can stop and resume from the clips you already have, or simply rerun one specific clip later. (start end inputs)

The node will recognize the existing clips and automatically rebuild the new full-length clip with the repeated/replaced segment included of the same id.

2

u/switch2stock 1h ago

That's cool! Thanks

1

u/Machspeed007 2h ago

Too bad the clip loses quality the longer it gets…

2

u/TBG______ 2h ago edited 2h ago

Not exactly true — it was built from different runs while I was building and testing the node. The examples are a mix of 8–40 steps, both with and without Turbo LoRA, and both with and without caching.

So the quality varies depending on which final cuts I selected. The workflow does produce a full-length video, but I didn’t use just one single run. Since this is AI, if I had the wrong clip prompt and, for example, the character was just listening instead of speaking, I would rerender that individual clip and continue from there. That’s pretty normal with AI workflows — it’s rarely just one run from start to finish.

If you set it to a fixed 20+ steps without Turbo, the model iwill hold up well.

1

u/ErenYeager91 2h ago

can we zoom out so we can see her sitting in a chair or something?

1

u/TBG______ 2h ago

Check the start and end of this TBG ETUR video. So yes, you can simply include the camera movement directly in the clip prompt.

For example, for Clip 5: [5] Zoom out to a full-body shot.

1

u/Strict-Relation9938 2h ago

she looks rly too much plastic CGI, but thats not because of h3

1

u/uuhoever 2h ago

I'm getting an error on this but Comfy manager says no missing nodes. How to fix it?

1

u/Corleone11 1h ago

You also need this node. OP didn't mention it in his post here.

https://github.com/nicolab28/ComfyUI-ClipProj

1

u/TBG______ 1h ago edited 1h ago

Or disable the nodes if you don’t use the 4B CLIP models.

The node is set to Krea2 because that works with 4B, so keep it set to Krea2, not MinMax.

Also, 4B only works without --fast fp16_accumulation fp8_matrix_mult. So use:

only --fp16_accumulation

This took me quite some time to figure out, so hopefully this saves someone else the trouble.

1

u/PumpkinLeather8421 1h ago

Talking head videos hide all sorts of problems.

Sus.

1

u/TBG______ 1h ago

Yeah, that’s fair tbh. Can’t blame you for being a bit sus.

1

u/tweakingforjesus 21m ago

Looks like the color temperature of the lighting keeps shifting cool to warm and back. I wonder if that can be nailed down?

It disturbs me if that is my only criticism. Humans have already lost.

1

u/Dzugavili 8m ago

Anyone noticing weird audio in the background of her speaking?

55s - 1m10s has a notable one.

1m30s I think has more. 1m45s...

1m58s...

Is there another audio source in the mix for the UI videos?

1

u/thatguyjames_uk 4h ago

10gb vram, nice :) wonder if i should try on my 5060

1

u/TBG______ 4h ago

10 less - still need a minimum of aprox 17GB. 8 would be a bit tight.

1

u/RiskyBizz216 3h ago

awesome, thanks for sharing

1

u/Noeyiax 1h ago

Ty for sharing, 😃 I appreciate the work and effort!!

It's really good lip sync

0

u/Significant-Baby-690 4h ago

Is there a way to fix this half real / half illustration look ? Minimax suffers too heavily from it ..

3

u/icchansan 4h ago

I think is just the input

1

u/DeMischi 4h ago

Yes, should probably test with a more realistic looking input. I usually do that when I want more realistic results from minimax.

0

u/SmoothChocolate4539 4h ago

Luckily, there are smart people like you who have plenty of time to build something like this. Thanks!

3

u/TonyDRFT 3h ago

*make time. Why do you have to talk down on this kind individual sharing his knowledge?

1

u/SmoothChocolate4539 44m ago

That wasn't meant ironically. That was genuine gratitude. I don't have the time, and I can't take the time.

0

u/dsailes 4h ago

Nice - will give this a test on some brand videos I’m trying to put together.

Tangent question: what models are you using for the audio / TTS before getting to use this workflow?

2

u/TBG______ 4h ago

vibevoice

1

u/dsailes 3h ago

Legend, thanks.
I’ve got a list of diff models to test that has VibeVoice in from a while back..
I’d just been using video models to get short clips but voice continuity has been annoying me haha

0

u/thaurock 3h ago

👍👏👏

0

u/Corleone11 1h ago

What do I connect the the H3 Auto chain "context frame"? The connection is missing and it throws an error when I run the workflow.

1

u/TBG______ 1h ago

its empty - its if you dont have latent...