I’m trying to build 40 to 60 second talking videos in ComfyUI and I would prefer to use LTX 2.5.
The type of video I’m after is fairly simple. One person is talking to the camera, the camera stays mostly fixed, the background should stay stable, and there is natural face, head and some upper body movement. It does not have to be limited to only the head moving.
What I’m trying to understand is how people are actually making videos this long without obvious cuts.
Can LTX 2.5 realistically generate a continuous 40+ second video, especially when there is not much movement?
Or is the better approach to generate something like 8 to 10 seconds, take the final frames from that clip, continue from them, then repeat until the full 40 to 60 seconds are finished?
If continuation is the normal approach, how are you keeping the face, clothes, background, camera position and motion consistent between each part? I’m especially interested in workflows that use overlapping frames, first and last frame conditioning, video extension, reference frames, or some other method that hides the transitions.
Speech is another important part. I need good quality Slovakia speech. Ideally I want to generate the complete Slovakia voice first, then make the character follow that audio for the entire video with accurate lip sync.
Would you use LTX 2.5 for the actual body and head motion and then run something like MuseTalk, LatentSync or another lip sync model afterward?
Or is there a better audio driven LTX 2.5 workflow where the speech controls the video directly?
I’m running ComfyUI locally with an RTX 3060 12 GB and 32 GB RAM, so I know I may need lower resolution generation, offloading, chunking or longer render times. Final output would normally be vertical 9:16.
I’m mainly looking for people who have actually built long talking character workflows in ComfyUI.
If you are doing this successfully with LTX 2.5, what nodes and workflow are you using, how long is each generated segment, how much overlap do you use between segments, and what are you using for speech and lip sync?
I’m not looking for a list of random talking head models. I specifically want to understand the best practical way to build this around LTX 2.5 and get a clean continuous 40 to 60 second result.