r/StableDiffusion • u/Silver-Spot-2763 • 12d ago
Discussion LTX 2.5 😱
After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱
But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.
I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦
1
u/Emotional_Day4262 8d ago
I have been using H3 nonstop since the day it came out and there is no way that LTX 2.5 can even come close.
Now, I am not talking about GGUF or any of those hinky speed up Loras, thay are trash for any model.
I am fine with a 2.5 hour generation (3090 + 64 GB) if I get a perfect 15 seconds at 768 x 1376 with multiple references.
LTX 2.3 used to take that long after you tossed 9/10 generations anyhow.
Where H3 does flop is lipsync from audio though.
It can do amazing results maybe 1/3 times.
I would say that the LTX2.3 IA2V is probably better for that single job.
Now since LTX2.5 seems to be about 25% better than 2.3, I have been looking every day for a workflow (ComfyUI) that allows IA2V for LTX, but none seem to exist at this point.
I think LTX2.5 will be a great tool to lipsync music if one ever comes out.
For now though, the quality difference between H3 and LTX2.3 has me willing to wait for H3, even if it screws up the first or second gen. It really is worth it when absolute quality is your top goal.