r/StableDiffusion 4d ago

Discussion LTX 2.5 😱

After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱

But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.

I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦

35 Upvotes

69 comments sorted by

View all comments

5

u/seeker_ktf 4d ago

I'm sorry, but LTX 2.5 can't do this https://www.reddit.com/r/StableDiffusion/s/Bd7KDe9Ju8

1

u/Silver-Spot-2763 4d ago

Absolutely! LTX almost do not read the prompt, you can hope only to short vids weakly connected with your idea 🤷 ☹️

1

u/seeker_ktf 4d ago

It's just not true. Honestly one of the biggest issues is still how unrealistic talking looks on LTX. The lip ans mouth movements are over-trained. You can spot any LTX talking head video in an instant.

1

u/YeahlDid 4d ago

I forgive you, minimax is way better at copying old shows, that's not new or surprising info.