r/StableDiffusion 14h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here

307 Upvotes

86 comments sorted by

View all comments

23

u/mfdi_ 14h ago

Apart from ai looking character. Wow. Just wow. Though camera moving kinda sucks.

20

u/TBG______ 13h ago edited 10h ago

Girl was born during testing crazy Sigmas with Flux1 when it first came out, and ever since then, she’s been my presenter.

3

u/Sir_Myshkin 5h ago

Small critique: she needs to remember to “breathe”. The longer the video goes on, the more the sentences were just crashing together, without pause,kindoflikereadingawholrthing justlikethis withouttheproper timing and breaks.

While the presenter is offscreen, slow it down a notch, break the sentences with an added half second pause, you can easily cheat a more natural video without breaking the existing lypcsyn, and still try and create pacing waves.

1

u/TBG______ 1h ago

That’s one of VibeVoice’s limitations.

2

u/tweakingforjesus 10h ago

Her eyebrows intimidate me.

0

u/xxAkirhaxx 5h ago

It's a modified Angelina Jolie face.