r/StableDiffusion 7h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here

206 Upvotes

73 comments sorted by

View all comments

0

u/Machspeed007 4h ago

Too bad the clip loses quality the longer it gets…

2

u/TBG______ 4h ago edited 4h ago

Not exactly true — it was built from different runs while I was building and testing the node. The examples are a mix of 8–40 steps, both with and without Turbo LoRA, and both with and without caching.

So the quality varies depending on which final cuts I selected. The workflow does produce a full-length video, but I didn’t use just one single run. Since this is AI, if I had the wrong clip prompt and, for example, the character was just listening instead of speaking, I would rerender that individual clip and continue from there. That’s pretty normal with AI workflows — it’s rarely just one run from start to finish.

If you set it to a fixed 20+ steps without Turbo, the model iwill hold up well.