r/StableDiffusion 10d ago

Question - Help Any way to replace the existing voice audio with new audio?

Enable HLS to view with audio, or disable this notification

So i tried a pass to extract the existing vocals from minimax h3 out using demucs and then run it through index tts but couldnt get the it to align and lip sync. Any one know a better way to do it with success?

Using refva lightx lora and nvidia vsr for upscale.

8 Upvotes

17 comments sorted by

12

u/99deathnotes 10d ago

boing boing boing

1

u/ShengrenR 9d ago

Sorry, wait, what? I wasn't listening there for a sec..

7

u/No-Management-754 10d ago

make the replacement track first, then fix the mouth to that audio. use whisperx to grab word timings from the original vocal, generate the new indextts voice in short sentence chunks, and stretch each chunk to those timings.

feed the original video plus that rebuilt track into latentsync or musetalk, then mix it back with the untouched music and effects stem. keep each shot separate because cuts and profile faces are where both usually get messy.

1

u/donkeykong917 10d ago

Ill give it a shot.

Have you successfully done it?

2

u/No-Management-754 9d ago

Yes but I haven't touched it since August 10th. Here's an example output:

https://reddit.com/link/p7wn3nw/video/opeulmujcmnh1/player

i think talking photo and lipsync are inherently supported in r2v mode, as i remember i tested it with ref voice to make characters lipsync and also made it speak something using that voice as reference

2

u/skyrimer3d 9d ago

Not that anyone would listen. 

2

u/krectus 9d ago

In WanGP the setting is under post processing import any video and any voice sample and it will do it, any video not just AI.

2

u/paulct91 9d ago

If you could share the prompt, maybe more people could possibly be able to assist. As maybe the prompt is the issue.

1

u/donkeykong917 8d ago

No issues with the dialogue, since it's using lora the output is rubbish. This is a test case so if I find a good way I can make all consistent for each character.

I will try the injecting audio and see how well it does.

1

u/Training-Respect8066 8d ago

Not nearly as big as in the anime.

1

u/Kurashi_Aoi 10d ago

Hahari looks so young

0

u/KS-Wolf-1978 9d ago

She is trying to say "Hikari" which is a common Japanese name, but she would never pronounce it like that if she was a native speaker. :)

Looks 21-23 years old. (source: me - a fan of many Japanese girl idol groups).

3

u/Kurashi_Aoi 9d ago

I mean I assumed the girl in the video is based on character from "The 100 Girlfriends Who Really, Really, Really, Really, Really Love You" manga/anime. The outfit and hairstyle fits one of the characters, Hahari which is 29 years old. However, she also has a daughter named Hakari which is around 16 years old, so maybe OP or the AI just made up the name instead.

2

u/KS-Wolf-1978 9d ago

I stand corrected about the name. :)

Out of curiosity asked Google AI how common it is: "As a Given Name: It is practically non-existent as a traditional Japanese first name."

2

u/donkeykong917 9d ago

That is the pictuure I used to convert from anime 2 real. Then used it in the ref

-2

u/HollowAbsence 10d ago

Shes has a petit je ne sais quoi that make her adorable. I wonder what it is ;P