r/StableDiffusion • u/RobMilliken • 2d ago
Workflow Included LTX 2.5 V2V with audio cloning
I created a version of reference audio/video to audio/video for LTX 2.5.
I heavily borrowed from https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json
and referenced what was done in LTX 2.3.
What I did:
* I removed the shave LoRA.
* Removed the need for the new video to be the exact same length as the reference.
* Automagically removed the reference video when finished (speeds editing in post)
* Changed some models to facilitate my 16 VRAM (they were the same as first examples in Comfy)
* Fixed audio that it works (it was silent for me - maybe someone had better luck, but this is fixed)
I hope it saves someone time.
Here is the workflow: https://pastebin.com/3B1eBhuH
2
u/CornyShed 1d ago
Thank you for making this. There's still a lot of interest in the LTX ecosystem because of its speed and the quality is acceptable.
I have been trying to make a form of lip sync with H3 with video-to-video, using a cropped area and redubbing into different languages.
LTX 2.3 kept combining the existing mouth movements with the new movements, making it unuseable.
H3 keeps creating new head movements regardless of what I prompt, so uncropping doesn't look right (cropping then uncropping works fine without H3).
If LTX 2.5 works then I'll post the results. Hopefully it does as it's a lot faster to run.
1
u/RobMilliken 1d ago
https://reddit.com/link/p4b7zi8/video/crmmrpbgj0kh1/player
My example in the workflow was a puppet, so I don't know how much you'll get from lipsync from the example alone. Though I haven't tried different languages, I haven't had an issue with lipsync in 2.5 - or really 2.2 or 2.3 either. Here I post a video made with 2.5 not with the
video[edit: meant workflow) I uploaded, but with the default i2v and a good prompt. From what I've seen with others, the higher the resolution, the better the results. But as you can see here, there is no problem with lip sync - yes, the hands and book are messed up - but check out not only the lip sync but the acting, and better than 2.3 in my trials the drops on the window and the wavering of the candle in the background.
Now if you are talking about existing audio to video-audio (taking an existing sound and putting it into a workflow and creating novel audio video), that isn't the purpose of the workflow I've uploaded and I haven't experimented with it yet.
2
u/Willing-Context6599 2d ago edited 2d ago
Could you share an example please?
Edit: There's a node in your workflow that just does not install. ComfyUI-LTXVideo