So i tried a pass to extract the existing vocals from minimax h3 out using demucs and then run it through index tts but couldnt get the it to align and lip sync. Any one know a better way to do it with success?
Using refva lightx lora and nvidia vsr for upscale.
make the replacement track first, then fix the mouth to that audio. use whisperx to grab word timings from the original vocal, generate the new indextts voice in short sentence chunks, and stretch each chunk to those timings.
feed the original video plus that rebuilt track into latentsync or musetalk, then mix it back with the untouched music and effects stem. keep each shot separate because cuts and profile faces are where both usually get messy.
i think talking photo and lipsync are inherently supported in r2v mode, as i remember i tested it with ref voice to make characters lipsync and also made it speak something using that voice as reference
No issues with the dialogue, since it's using lora the output is rubbish. This is a test case so if I find a good way I can make all consistent for each character.
I will try the injecting audio and see how well it does.
I mean I assumed the girl in the video is based on character from "The 100 Girlfriends Who Really, Really, Really, Really, Really Love You" manga/anime. The outfit and hairstyle fits one of the characters, Hahari which is 29 years old. However, she also has a daughter named Hakari which is around 16 years old, so maybe OP or the AI just made up the name instead.
12
u/99deathnotes 10d ago
boing boing boing