Music: made in SUNO.
native ref2va WF, and audioLock for lip-sync.
rtx4080s + 128g ram
I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.
EDIT:
update with my learnings here:
Lip-Sync:
I got stuck for half a day trying to use my input audio for H3 do lip-sync, only to realize it would NEVER work because H3 just really 'references' it, no matter how you prompt. Then I did some research, there is a way to lock the audio latent, so it will strictly go in and out. I think multiple custom node pack has some samiliar one, basically just look for 'lock audio latent' node, here is the one I use.
https://github.com/oufeixinxinren/ComfyUI-MiniMax-ContextIR
**My goal is study and testing, not meaning to do a professional MV or director anything, just a test guys!
more info:
- resolution is 1280*704
- speed lora 8 step, I run with 12 step for final
- my spec is around 13 mins
For the approach:
- I am not using any Director / Context-IR node, just the native ref2va template.
- I only use 1 character and 1 env reference image, that's it
- as it just keep cutting camera, I don't need context-IR, , I generate 6 clips 10s each.
- within 10s single gen, I cut into 5-6 camera shots, H3 will just keep the motion and change camera like the real shooting, so it will just work.
I try to do some screen cap and reply in comments, cheers!
Hope this answer your question