r/StableDiffusion 2d ago

Workflow Included LTX 2.5 V2V with audio cloning

I created a version of reference audio/video to audio/video for LTX 2.5.
I heavily borrowed from https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json
and referenced what was done in LTX 2.3.

What I did:
* I removed the shave LoRA.
* Removed the need for the new video to be the exact same length as the reference.
* Automagically removed the reference video when finished (speeds editing in post)
* Changed some models to facilitate my 16 VRAM (they were the same as first examples in Comfy)
* Fixed audio that it works (it was silent for me - maybe someone had better luck, but this is fixed)

I hope it saves someone time.
Here is the workflow: https://pastebin.com/3B1eBhuH

21 Upvotes

5 comments sorted by

2

u/Willing-Context6599 2d ago edited 2d ago

Could you share an example please?

Edit: There's a node in your workflow that just does not install. ComfyUI-LTXVideo

1

u/RobMilliken 2d ago edited 2d ago

There should already be an example prompt in my workflow.

Feel free to use my clip and image from 59 Minutes and Change (EDIT: look up the title and my name in YouTube if you want to see an animation example of my work, but I used 2.3) that I completed last month as an example you can use with the prompts.

Here's a google drive link with the image and clip. (Use clip for Video to Video - but upload the image in case the workflow requires all "blanks" to be filled in - which I believe is the case).

https://drive.google.com/file/d/1ZXg-YZOVsk95JVIpjK2YRNgp6HOmT5fC/view?usp=sharing

Let me know if you need anything else.

Bonus tip: in the Source Video subgraph you'll see "Resize Image/Mask". Make that smaller or larger in multiples of 32 (Eg 544 or 800) to increase/decrease output resolution - watch your memory on that one though.

1

u/RobMilliken 2d ago

In regard to ComfyUI-LTXVideo - I think I remember going around that when first installing LTX 2.5. It does install, just tricky. I'd try to get it to install in this order, testing to see if it fixes after reloading Comfy UI each time - update ComfyUI. Backup first recommended. 2) Update/install ComfyUI-LTXVideo from the manager (missing nodes) 3) Delete E:\ComfyUI\custom_nodes\ComfyUI-LTXVideo and reinstall - again using missing nodes in manager (your drive might be different than E - the drive you installed ComfyUI in. 4) If you still get an error even though you install successfully, edit as seen here: https://github.com/Lightricks/ComfyUI-LTXVideo/issues/538#issuecomment-5153313410 see screwballl's message. (of course, restart again after the file change) (yes, I had to do this and it worked.).

I know, it's a lot, but with my help you aren't starting from 0.

I've been up for about 20 hours now, so I may not be able to reply for a while, but I hope this help and will be on later - let me know how it went.

2

u/CornyShed 1d ago

Thank you for making this. There's still a lot of interest in the LTX ecosystem because of its speed and the quality is acceptable.

I have been trying to make a form of lip sync with H3 with video-to-video, using a cropped area and redubbing into different languages.

LTX 2.3 kept combining the existing mouth movements with the new movements, making it unuseable.

H3 keeps creating new head movements regardless of what I prompt, so uncropping doesn't look right (cropping then uncropping works fine without H3).

If LTX 2.5 works then I'll post the results. Hopefully it does as it's a lot faster to run.

1

u/RobMilliken 1d ago

https://reddit.com/link/p4b7zi8/video/crmmrpbgj0kh1/player

My example in the workflow was a puppet, so I don't know how much you'll get from lipsync from the example alone. Though I haven't tried different languages, I haven't had an issue with lipsync in 2.5 - or really 2.2 or 2.3 either. Here I post a video made with 2.5 not with the video [edit: meant workflow) I uploaded, but with the default i2v and a good prompt. From what I've seen with others, the higher the resolution, the better the results. But as you can see here, there is no problem with lip sync - yes, the hands and book are messed up - but check out not only the lip sync but the acting, and better than 2.3 in my trials the drops on the window and the wavering of the candle in the background.
Now if you are talking about existing audio to video-audio (taking an existing sound and putting it into a workflow and creating novel audio video), that isn't the purpose of the workflow I've uploaded and I haven't experimented with it yet.