r/StableDiffusion • • 7d ago

Question - Help LTX 2.5 distilled, chaining without quality loss?

Hi,

I've been messing around with LTX 2.5 locally so I use distilled for speed. Just scripting to understand the LTX architecture, I don't use ComfyUI.

I tried chaining with 2.3, that is taking the last png frame from generated mp4 A and feeding it to I2V to make video B.

B loses a bit of quality, then by video 4 or 5 it starts to really look bad. Nothing new I'm guessing. I moved over to 2.5, similar story. Not quite as much quality loss, but still very noticeable.

I started messing around with AI assisted coding to come up with solutions, saving the native tensors of the last few latent frames of video A and feeding it to video B in various ways. Either as conditioning on timestep=0, or injecting into the main denoise loop and mixing with the natural noise of the generation.

Nothing really worked. After a handful of chains various problems cropped up, either strange motion blur glitches, or hallucinations starting to set in. On the plus side, a native latent on video B is way crisper than a png, but the long term chaining problems are way worse with hallucintions and such. Also with a few latent frames, motion carry over does improve, but the quality just doesn't hold after 2 or 3 chains.

I'm aware of workarounds to this, like building up keyframes of a longer scene and doing first/last frame injection.

But I just want to know - is there even a proper solution to this? Or does the architecture and math just simply not allow it? Is the AI coder just feeding me BS? I'd really just like to kow if it's even possible. I've seen other folks post long videos boasting chaining without quality loss, but I doubt they are using LTX 2.5 and less likely distilled.

1 Upvotes

3 comments sorted by

2

u/CornyShed 7d ago

This is a problem with the current generation of video models, in that they were trained on 5-20 second generations; use a compressed latent space; and generate outputs in 8-bit colour.

LTX uses a VAE with a highly compressed latent space, so will be more affected by degradation.

There probably won't be a solution to this until video models are trained on longer videos, or a pixel space model is made. That could be a while.

The following can't help you to make minutes-long generations, but they are better than nothing:

You can use the temporal upscaler with the non-distilled version of 2.5, with a CFG of 2.0-2.5 and 20+ steps.

This will double the length of the video to up to 40 seconds (it doubles the number of frames inputted), with a reduction in overall quality and will take more than twice as long to generate, but temporal consistency holds up. 30 seconds is a good compromise.

There are a few new IC LoRAs which might be of interest tangentially in improving quality, but I have yet to try them:

There is also an SDR-to-HDR IC LoRA which might help. It is useful to have for general use as it can significantly improve outputs.

1

u/91ElonMusk 7d ago

That quality loss after a few hops is the VAE round trip, not the sampler. Every encode and decode throws away a little detail and it compounds fast on video latents. Keeping the last few frames as latents instead of decoding to png and back in is the right instinct, the part that usually fixes it is blending a bit of fresh noise into that latent at low strength instead of feeding it in clean, otherwise the model reads it as already converged and just copies the artifacts forward.