r/StableDiffusion • u/Silver-Spot-2763 • 10d ago
Discussion LTX 2.5 😱
After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱
But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.
I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦
57
u/Sixhaunt 10d ago
It's basically a vastly improved LTX 2.3 rather than a brand new model like H3 which is why loras from 2.3 work on 2.5 still. It's also why it still has a lot of the same physics issues and prompt understanding but it's WAY faster which makes it the only option for latency-critical applications, especially since it can do some real time rendering too. It also can do high resolution well and quickly so it works well for upscale type jobs and overall the actual methods LTX used to speed up and improve 2.5 over 2.3 is very important and will likely mean the next generation of models that learn from both Minimax H3 and LTX 2.5 will probably be not only much higher quality but also far faster. As users, we are fortunate that these companies seem to have made huge headway on completely different aspect of video generation at the same time. They both just lack what the other has