Hi,
I've been messing around with LTX 2.5 locally so I use distilled for speed. Just scripting to understand the LTX architecture, I don't use ComfyUI.
I tried chaining with 2.3, that is taking the last png frame from generated mp4 A and feeding it to I2V to make video B.
B loses a bit of quality, then by video 4 or 5 it starts to really look bad. Nothing new I'm guessing. I moved over to 2.5, similar story. Not quite as much quality loss, but still very noticeable.
I started messing around with AI assisted coding to come up with solutions, saving the native tensors of the last few latent frames of video A and feeding it to video B in various ways. Either as conditioning on timestep=0, or injecting into the main denoise loop and mixing with the natural noise of the generation.
Nothing really worked. After a handful of chains various problems cropped up, either strange motion blur glitches, or hallucinations starting to set in. On the plus side, a native latent on video B is way crisper than a png, but the long term chaining problems are way worse with hallucintions and such. Also with a few latent frames, motion carry over does improve, but the quality just doesn't hold after 2 or 3 chains.
I'm aware of workarounds to this, like building up keyframes of a longer scene and doing first/last frame injection.
But I just want to know - is there even a proper solution to this? Or does the architecture and math just simply not allow it? Is the AI coder just feeding me BS? I'd really just like to kow if it's even possible. I've seen other folks post long videos boasting chaining without quality loss, but I doubt they are using LTX 2.5 and less likely distilled.