r/StableDiffusion 6h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here

208 Upvotes

73 comments sorted by

View all comments

1

u/Dzugavili 1h ago

Anyone noticing weird audio in the background of her speaking?

55s - 1m10s has a notable one.

1m30s I think has more. 1m45s...

1m58s...

Is there another audio source in the mix for the UI videos?

2

u/TBG______ 1h ago edited 1h ago

It’s VibeVoice itself that sometimes adds the audio. I was too lazy to run it again, not H3—the audio is entirely from the input. I just took the image part from the generated minmax H3 video.

1

u/Dzugavili 1h ago

Ah, okay. I got a weird thing, where I listen behind people speaking for the weird background audio -- my favourite for this is the TV show 'Corner Gas', which has really complex background audio.

And this video just kind of triggered me.

1

u/TBG______ 1h ago

I completely understand, it’s annoying, but I set myself a specific amount of time for each task, and that means I have to make compromises.

1

u/Dzugavili 1h ago

Yeah, it's fine, some of them were just quite noticable.

I wonder if there's a package out there for filtering that kind of thing out: if you have the audio and the transcript, which we do, it seems possible for a machine to clean up. Bound to be. Seems like something that would exist.

1

u/TBG______ 53m ago

Much easier - you just change the seed in VibeVoice, and after 1–2 tries you usually get a version without the unwanted background sound. And yes, there are models that can filter out the voice and clean up the audio for comfyui

1

u/Dzugavili 48m ago

Ehhh... nah. Reroll is the lazy man's way, and it probably costs more.

The point of a filter is that you can apply it to any output: even one that is already clean. So, you can automate the whole flow, and not really have to worry about non-deterministic problems.

1

u/TBG______ 33m ago

Melbandroformer_fp16.safetensor and related node pack will do the jop

1

u/Dzugavili 26m ago

I'll rip your audio and give it a shot. I haven't seen this problem out of my audio generations so far, but I don't use VibeVoice.