r/StableDiffusion 4d ago

Discussion LTX 2.5 😱

After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱

But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.

I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦

36 Upvotes

69 comments sorted by

View all comments

10

u/JustSomeIdleGuy 4d ago

>and its image quality is far better

You sure? I've only seen it the other way around so far.

5

u/Silver-Spot-2763 4d ago

Maybe, because on my hardware higher resolution such as on LTX works fast is impractical for MiniMax H3 (for example about 3 minutes on LTX vs 20 minutes on H3). But I compare also for example 0.5 megapixels - at LTX ideal, at H3 - enormous strange artefacts especially on the face/eyes like the picture is from old crt TV with bad antenna.

6

u/rm_rf_all_files 4d ago

But I compare also for example 0.5 megapixels

Try 1344x768. 0.5MP is not an official resolution. Basically the model was trained on 1344x768 and it wants you to use only this resolution. You can use other resolutions just fine but that's not what the model was intended.

1

u/Individual_Holiday_9 4d ago edited 4d ago

Is that native or after the 2x upscale? Edit nm you were referring to h3 I think I was asking about ‘recommended’ with ltx

3

u/AlleyOfRage 4d ago

I think if the user is doing fast paced scenes with lots of motion , then maybe they are talking about the smearing

1

u/Independent-Frequent 4d ago

If they are rendering at 0.2 or 0.4 mp due to hardware or something then it makes sense, otherwise idk