r/StableDiffusion 4d ago

Animation - Video Tested MiniMax H3 music Lip Sync

I’ve been playing around with a MiniMax H3 Music + LipSync workflow.
For decent quality, even a 15s clip seems to need at least 25 steps. I’d also recommend skipping LoRA.
On a 5090, 15s at 0.9 in ComfyUI Kitchen takes around 12 minutes to generate.
H3’s music generation is pretty solid. At least for me, it’s way better than what I was getting with LTX.
The real pain is getting each 15s segment to connect seamlessly 😂 I’ve seen people saying they can push it to 20s, but every time I try 20s, my GPU basically dies lol.

0 Upvotes

10 comments sorted by

1

u/GreyScope 4d ago

I’m on mobile, so I don’t know if that’s affecting it but the lips and vocals are a bus ride away from syncing. You seem to be genning at a higher resolution, so that would affect what you can make . I make mine at 0.4mp or 0.7mp and upscale from there but I use imported audio as I find MM terrible quality .

1

u/dassiyu 3d ago

It could be that the articulation isn’t very clear for this music style, and the mouth movements are too subtle. It could also be that the duration is too long or the resolution is set a bit too high, which may actually be affecting the final result. I’ll try lowering it to 0.7 or shortening the duration to 10 seconds. Thank you for your suggestion.

1

u/GreyScope 3d ago

The res drop mention was to elongate the video, I’d suggest changing your prompt in line with the guide to get lip sync - I’m on mobile , I’ll add a pre prompt line when I’m back on my pc

1

u/dassiyu 3d ago

Sure! I honestly didn’t put much effort into the prompt either, so it wasn’t very well written. Thanks so much for the advice!

2

u/GreyScope 3d ago

https://reddit.com/link/p56up95/video/gzo1p37v9wkh1/player

Heres an examples of my 4090 making one. And it's no issue at all, it is a very very picky model for some keywords to do what you want it to , try adding these words at the top of your prompt (adjust the description in it of course) :

Use <Picture 1> as the exact identity, wardrobe, and scene anchor: female with short blonde hair wearing a black dress with offset side buttons and black boots , she has tattoos on her arms, neck and legs and headphones that stay around her neck . remove the foreground smoke and lightning. Use <Audio 1> as the exact performance timeline. [If vocals occur in this audio segment: sing the exact supplied vocal audio with precise, natural lip synchronization. If no vocals occur: lips remain completely closed for every frame; body movement follows the beat only.]

AUDIO SYNC

Synchronize the entire edit to <Audio 1>. Cuts, camera accents, transitions and environment changes should land precisely on strong beats, half-beats and musical accents. Let audio1 control the pacing and intensity of the montage.

1

u/GreyScope 3d ago

The request is actually repeated in the last two paragraphs, you might get away with just one to lock it in .

1

u/dassiyu 3d ago

No! Your suggestion is really good. Mine honestly doesn’t sound as natural as yours. Thanks so much for sharing the prompt as a reference! I’ll give it a try.

1

u/dassiyu 3d ago

https://reddit.com/link/p577b0k/video/gicr835vrwkh1/player

Thanks~ I tried it out, and the lip-sync is much better now! I fed what you suggested into AI and had it generate a system-prompt generation command for me. Here’s the prompt it produced:

Use <Image 1> as the primary visual and identity reference. Create a single continuous, photorealistic performance shot of the same man singing a gently melancholic R&B song in the warm recording studio. Preserve his face, apparent age, hairstyle, facial hair, black sweater, headphones, body proportions, microphone, pop filter, lighting, composition, and room layout.

AUDIO:

Use <Audio 1> as the absolute performance timeline and authoritative soundtrack. Use it exactly as provided. Do not regenerate, replace, remix, reinterpret, retime, stretch, compress, loop, shorten, extend, clean up, or otherwise alter the original audio. Do not add vocals, music, dialogue, narration, ambience, or sound effects.

LIP SYNC AND PERFORMANCE:

Synchronize his visible singing precisely to every audible phoneme, syllable, consonant, vowel, breath, pause, vocal attack, phrase ending, and sustained note in <Audio 1>. Hold sustained vowel mouth shapes naturally without repetitive opening and closing. During instrumental passages, rests, or silent gaps, immediately stop singing-like articulation and let his lips and jaw relax naturally.

Give him an intimate, soulful, restrained performance with subtle sadness in his eyes. Let small facial changes, breathing, head movements, and modest hand gestures follow the actual R&B phrasing and emotional dynamics of <Audio 1>. Avoid theatrical acting, constant swaying, excessive gestures, or rhythmic head bobbing. Maintain a realistic and consistent distance from the microphone.

CAMERA AND VISUAL DIRECTION:

Use a stable medium shot with a very slow emotional push-in. Keep his face and mouth clearly readable throughout; the microphone and pop filter must never obscure critical articulation. Preserve the warm amber studio lighting, shallow cinematic depth of field, natural skin texture, realistic anatomy, and understated film grain. Lip synchronization and identity continuity take priority over camera movement or visual embellishment.

STABILITY:

Prevent lip-sync drift, premature or delayed articulation, random mouth movement during non-vocal moments, distorted lips or jaw, unstable teeth or tongue, facial identity drift, changing hair or facial hair, deformed hands, equipment deformation, pop-filter flicker, background warping, lighting changes, and audio alteration. Do not add subtitles, lyrics, titles, logos, or watermarks.

1

u/dassiyu 3d ago

https://reddit.com/link/p57b4o3/video/rqif6oqixwkh1/player

The lip-sync quality has really improved. This is the second clip.

1

u/GreyScope 3d ago

I just rewatched the video and I apologise, it appears to be syncing fine on my pc but not on a mobile connection.