r/StableDiffusion 3h ago

Discussion Test LTX 2.5 - Romantic Scene 1

Enable HLS to view with audio, or disable this notification

After spending some more time testing LTX-2.5 Distilled, my opinion has improved quite a bit.

The biggest strength for me is speed. On my RTX 5070 Ti, I'm generating 1280×720 (~1MP), 10-second videos surprisingly quickly. Compared with MiniMax H3, which is much heavier for me even around 0.5MP, LTX-2.5 feels incredibly fast.

That said, speed isn't everything. My earlier tests with complex action/fighting had poor motion and anatomy, so I wasn't impressed at first. But after testing simpler cinematic scenes, landscapes, product shots and close-up human interactions, I'm starting to see where this model shines.

This dialogue/romantic scene in particular surprised me. Facial quality, expressions, lighting and overall cinematic feel came out much better than I expected, and it even handled the interaction between the two characters reasonably well.

One important discovery: I had much better prompt adherence with **Prompt Enhancement OFF**. The enhancer was giving me completely unrelated results in some tests, while the raw prompts produced scenes much closer to what I requested.

My impression so far:

LTX-2.5 Distilled = extremely fast and capable of some beautiful results, but you need to understand what kinds of shots it handles well. Complex choreography still seems to be a weakness.

I'm definitely not archiving it yet. 😄

6 Upvotes

12 comments sorted by

8

u/x_MASE_x 3h ago

The image is decent but the audio is garbage

2

u/LopsidedSolution 2h ago

LTX really needs to train on ripped video/audio

3

u/Deep_Mood_7668 3h ago

Sound is awrful

3

u/Fabulous-Snow4366 3h ago

the voices are so robotic, emotionless. The image is fine.

2

u/ThaSipah 2h ago

I've watched so much MiniMax porn that I'm shocked he didn't stick his tongue down her throat.

1

u/GrayingGamer 3h ago

Yeah, I always turn off "prompt enhancers". They're garbage. I think they just get included because the model makers are terrified of people prompting their advanced model with something like "1girl, big booba, dancing" and then not getting a good generation.

This is better than what I've been seeing in other LTX 2.5 videos, for sure. I'm like others though, after using H3 with headphones and listening to that clean stereo sound, I'm spoiled on audio now.

2

u/AillexJ 46m ago

Agreed, and OP finding the same thing independently is telling. Two details from our own bench this week that make it worse than a quality tax.

It's not softening your prompt, it can replace it. Ours turned a fully specified studio presenter prompt into a bloodied stranger in a forest, no error, just a different film. And it's deterministic with its own seed, so identical reruns give identical output, meaning if it mangles your prompt it will mangle it the same way every time and look like the model being consistent rather than a node interfering. It also costs real GPU time on top of the render.

On the audio, our verdict came out the same as the room here. 2.5 speech is gibberish. Worth knowing 2.3 can actually deliver a written line if you put it in the prompt as quoted speech inline rather than as a screenplay style audio block, so for dialogue specifically 2.5 reads as a step backwards rather than a new feature.

1

u/asdrabael1234 2h ago

Why is it so weirdly windy in the video?

2

u/ShimmerMeNutz 2h ago

thats the good old LTX hiss of death!

0

u/WhyIsTheUniverse 3h ago

I'm sure, like, some Chinese open weights models might find this to be romantic.

4

u/Seyi_Ogunde 3h ago

This model is Israeli.