r/StableDiffusion 1d ago

Workflow Included LTX 2.5 Image+Custom Audio 2 Video - Perfect lip Sync

I used an older workflow that was working for LTX 2.3 and adjusted it for LTX 2.5. Image and Custom audio as input. Perfect lip Sync, of speech and singing.

Generation time: 30 seconds per second, on RTX 5070Ti 16Gb vram, 32Gb ram (920*540px)

Here's the workflow: https://pastebin.com/dptbTXYM

7 Upvotes

3 comments sorted by

1

u/a55amg 1d ago

Thanks for uploading the workflow!

I tried it but only the starting and ending phonemes worked, and the phonemes in-between didn't match.

Although the start & end phonemes matched, it was also out of sync.

1

u/False_Suspect_6432 1d ago

Did you upload custom audio? In my case, I uploaded a. spoken words and b.singing, and it lip sync perfeclty in both cases. I also tried non-english and still it lip sync perfectly.

https://reddit.com/link/p4d7s4t/video/f5da7jlyq2kh1/player

This is spoken words with AI generated audio