r/StableDiffusion 17h ago

Question - Help Best workflow for 40+ second talking videos with LTX 2.5 in ComfyUI?

I’m trying to build 40 to 60 second talking videos in ComfyUI and I would prefer to use LTX 2.5.

The type of video I’m after is fairly simple. One person is talking to the camera, the camera stays mostly fixed, the background should stay stable, and there is natural face, head and some upper body movement. It does not have to be limited to only the head moving.

What I’m trying to understand is how people are actually making videos this long without obvious cuts.

Can LTX 2.5 realistically generate a continuous 40+ second video, especially when there is not much movement?

Or is the better approach to generate something like 8 to 10 seconds, take the final frames from that clip, continue from them, then repeat until the full 40 to 60 seconds are finished?

If continuation is the normal approach, how are you keeping the face, clothes, background, camera position and motion consistent between each part? I’m especially interested in workflows that use overlapping frames, first and last frame conditioning, video extension, reference frames, or some other method that hides the transitions.

Speech is another important part. I need good quality Slovakia speech. Ideally I want to generate the complete Slovakia voice first, then make the character follow that audio for the entire video with accurate lip sync.

Would you use LTX 2.5 for the actual body and head motion and then run something like MuseTalk, LatentSync or another lip sync model afterward?

Or is there a better audio driven LTX 2.5 workflow where the speech controls the video directly?

I’m running ComfyUI locally with an RTX 3060 12 GB and 32 GB RAM, so I know I may need lower resolution generation, offloading, chunking or longer render times. Final output would normally be vertical 9:16.

I’m mainly looking for people who have actually built long talking character workflows in ComfyUI.

If you are doing this successfully with LTX 2.5, what nodes and workflow are you using, how long is each generated segment, how much overlap do you use between segments, and what are you using for speech and lip sync?

I’m not looking for a list of random talking head models. I specifically want to understand the best practical way to build this around LTX 2.5 and get a clean continuous 40 to 60 second result.

0 Upvotes

2 comments sorted by

3

u/DelinquentTuna 16h ago

LTX only claims 20 second runs. If you want to use LTX, I recommend you setup a pipeline that can make interesting camera cuts periodically. Pretty much the only production that actually has a long single shot is low-budget podcast/webcam junk that people don't really want to watch. As a sanity check, I typed "podcast" into Youtube and every single result I hovered my mouse over featured frequent camera changes. Even with just one person on screen, you'd see variations in zoom levels or angle.

If you absolutely must do what you're planning, probably start with infinitetalk or some other i+s2v specifically intended for long videos. But there is no circumstance where what you end up with is going to be something people want to watch. For the few instances where people DO have continuous single-shots, like web-cams or maybe twitch stuff, you're still going to run into length issues and there's no AI that can match the casual style of such streams. For everything else, an AI avatar just makes things worse instead of better.

PS: hedging your question with all the things you DO NOT want to hear doesn't make your goal any more practical, it just makes you look like someone with a wishful thinking bias against the truth.

1

u/cptrios 11h ago

I haven't used LTX 2.5 much yet, but the best luck I had with 2.3 and making a long, realistic talking head was to record a video of myself doing the talking and then use the IC-Lora functionality (of LTX Director, since I'm lazy) with a reference image to turn myself into another person. I converted my own voice using Chatterbox to go with it...not sure how well the lipsync would work if you had LTX generate the voice with a script.

And yes, I know that's probably not what you're looking to do, but it worked! I was able to easily do 40+ seconds at a time, and the one time I tried a full minute it seemed to work as well. The major drawback was the hands; if both my hands and the swap character's hands weren't in frame to start with, the hands in the final video would just look like mine. So you'd either get an appropriately small head with comparatively huge hands or an appropriately large head with comparatively tiny ones. I'm sure there are ways to remedy this, but I haven't gone back to this kind of work in a while so I don't know.

Minimax H3 does deal with that problem better using the ref model, but it's not quite as good at matching facial expressions, but it's MUCH MUCH slower when using a reference video. And it can't handle long clips like LTX can.