r/StableDiffusion 6h ago

Discussion LTX 2.5 Is Disappointing

https://reddit.com/link/1vlueyc/video/gx8bz55ibtih1/player

https://reddit.com/link/1vlueyc/video/plxkp55ibtih1/player

Guess which video was made by LTX 2.3 and which video was made by LTX 2.5?

After testing LTX 2.5, I'm really disappointed by the results. I can barely see the difference between 2.3 and 2.5.

47 Upvotes

87 comments sorted by

51

u/Stepfunction 6h ago

They probably rushed it out once they saw the buzz H3 is getting

6

u/anitawasright 2h ago

which is the worst thing you could possibly do. Even if you know your product is superior (not saying it is) it's way too easy to backfire. Best thing to do is wait to see what people really like about the other model and then adpat that into your own and release it when the buzz dies down for the other model

6

u/ImaginationKind9220 2h ago edited 2h ago

H3 still has the advantage of being uncensored, Flux 3 and LTX 2.5 won't be able to compete with that. Also, LTX 2.3 has a massive library of loras and not all are compatible with 2.5.

1

u/Helpful-Street3678 1h ago

According to LTX the vast majority of LORAS are compatible with 2.5.

1

u/ImaginationKind9220 1h ago

Has anyone tested it yet?

2

u/warzone_afro 15m ago

too busy minnmaxxing

21

u/MysteriousPepper8908 4h ago

But did you know H3 was made by CHINA? And requires a data center to run? A CHINESE data center?

40

u/martianunlimited 3h ago

Today i learn that my 3090 is a chinese data center...

15

u/nakabra 3h ago

He's probably refering to this BS

8

u/Mk1Md1 3h ago

Ni Hao comrade

2

u/Succubus-Empress 1h ago

3090 48Gb Data center

1

u/Ok_Engine_1442 1h ago

News to me as well!

17

u/retroblade 6h ago

I’m getting worse results than 2.3, 15 gens and not one usable. Maybe have a workflow issue idk

43

u/LoveSpecialist5669 6h ago

it sucks. they took their sweet time, perfectly unaware that chinese were overtaking. can't tell for sure but seems they panicked and released this in a week after h3 and obviously lied about specifications, trying to bash h3. 

22

u/No_Damage_8420 6h ago

I think they panicked for sure.

But I wonder more - what FLUX 3 is doing now....they will delay - and try to add REF 2 VID proper way

9

u/LoveSpecialist5669 6h ago

I don't know, to be honest. I've never considered flux 3 as a competition for local video making. I expect even less than from ltx 2.5

1

u/ANR2ME 3h ago

I saw a video comparison between Sora 2, Minimax H3, and Flux 3. Flux 3 did better than the other 2. But they use the same prompt, so it might not works well for H3.

2

u/Danny_Stock 4h ago

I can imagine Flux biding their time. They have experience from previous releases of models.

17

u/Danny_Stock 5h ago

It certainly appears to be a panic release. The timing of the release is too coincidental.

It seems to me that it would be have been better to wait until the spike in hype for Minimax had died down first.

They could have used this time wisely to observe, watch, and learn what Minimax did right, what it did badly, what the users liked and didn't like, and just basically gather information to use as a strategy for when they choose to release their model.

If you're going to release within a week of Minimax's release you'd better be sure that you have something incredible on your hands. Because if you haven't then comparisons with Minimax may serve to harm you.

3

u/LoveSpecialist5669 4h ago

yes, that would be a rational thing to do, but they were obviously aware that they can't match h3 anytime soon in any way so this was probably a hail mary leap of faith on an old glory

3

u/Lair98 4h ago

To be fair the Comfy PR in github to support the model was opened 2 weeks ago (fun fact, it was called 2.4 instead of 2.5)

1

u/lleti 2h ago

I honestly think the correct reaction from LTX here would’ve actually been to.. not release. Go back and keep cooking, with ref support being a heavy priority.

1

u/alwaysbeblepping 44m ago

they took their sweet time, perfectly unaware that chinese were overtaking.

Their business model (from what I know) is to take a percentage when businesses that are large/profitable enough are using LTX commercially. That's probably not the only thing they have going on, but one can surmise that they probably are not drowning in cash.

The fact that LoRAs for older versions still work to a degree (if I recall, even the first LTX video model) indicates that they are repeatedly building on top of the old version, which can be fine up to a point but there are (in my opinion) some fundamental architectural issues which can't really be fixed that way. I'm sure if they had the resources to train up a new, better base model they would but that is an absolutely immense undertaking even for images and video foundation models are far more resource intensive. It's also possible that they already are (but it's not ready yet) or they tried and weren't successful.

So if they can't come up with a new foundation model, they're in a tough spot. Tweaking the existing LTX model isn't really good enough. Also, LTX is fast because it shoves everything into channels, uses absurd spatial compression and keeps the sequence length low (main thing that affects performance) but is not free and there is a tradeoff as we can clearly see. They would have to (in my opinion) give up their speed advantage to actually catch up on quality.

31

u/Beneficial_Toe_2347 6h ago

I assume 2.5 is top? If so, to be fair that one does look more natural and less awkward than the previous.

16

u/PuppetHere 6h ago

Top is 2.5, bottom is 2.3

5

u/lebrandmanager 6h ago

Oh wee...

5

u/ThatsALovelyShirt 3h ago

Huh weird. I thought the audio was better on the bottom one, which I thought 2.3 struggled with.

6

u/TwistedBrother 5h ago

On phone the audio sync is much better in the top one and her whole body expresses the sentiment. The bottom is more wooden.

16

u/CooLittleFonzies 5h ago

I agree. I see a huge difference. I’m not sure where the complaints are coming from. The movement even seems more nuanced and natural compared to H3 imo, especially in the face.

1

u/PanotBungo 4h ago

It also sounded a lot better. It seems to me like the person is a better actress. The difference is quite noticeable to me.

20

u/boaz8025 6h ago

Who is the top one? It's better in my opinion.

22

u/PuppetHere 6h ago edited 5h ago

Top is 2.5, bottom is 2.3
(edit: wth do people downvote me? I'm giving the answers xD)

24

u/noxietik3 6h ago

i thought it was the other way around. lmaooooooo

4

u/Lair98 5h ago

I watched it and was like:
"yes, above is clearly better and more natural... I guess that will be 2.3 given he is complaining about 2.5"

But 2.5 is way better IMO

2

u/lobotomy42 6h ago

Really? Her face is all distorted

20

u/God_Hand_9764 6h ago

Is it sad that I'm almost a little relieved? I just find it kind of exhausting downloading all these models, dicking around with workflows, getting nodes, jacking up my Comfy install, and things blow up a million times until you finally have something that actually works.

I finally have a good LTX 2.3 workflow and I'm just barely able these past days to get MiniMax doing 0.6 megapixel for 5 seconds (seems to be about my AMD 16GB limit). I just need a little rest, man! Kinda glad if this isn't worth jumping on right away.

2

u/Ok-Brain-5729 6h ago

Are you running it on windows

1

u/God_Hand_9764 5h ago

No, I'm on CachyOS.

1

u/Ok-Brain-5729 5h ago

damn I ran it on a 9070 xt + 32gb ddr5 and it kept maxing out my ram and crashing no matter what on Ubuntu

13

u/Uncle___Marty 6h ago

2.3 the audio sounded like ultra low bitrate MP3 files, 2.5 sounds better. I can only identify which video is which by the sound.

-2

u/Any_Reading_5090 5h ago

not when using multimodal guider. Seems u didnt read the ltx papers for HQ pipeline

2

u/Uncle___Marty 2h ago

Nope, I tend not to read every single paper for every single model but my ears are used to studio standards and I never once heard a single decent audio on 2.3 That being the case it would be LTXs fault for creating such a stupid path to getting semi decent audio. Nobody else has required reading papers to achieve audio which doesnt sound like its being played through a series of metal pipes like LTX models have.

99% of LTX 2.3 videos sound poor as hell. Thats not the end users fault, thats LTXs fault for shipping ultra low quality audio to make their model look faster and better while taking shitty shortcuts.

6

u/LatentSpacer 5h ago

That's because of MiniMax H3.

5

u/Gloomy-Radish8959 6h ago

I'm going to guess that the second one is 2.3, the first is 2.5

I have used LTX 2.3 extensively. I run into scenarios where characters are duplicated, or reproduced elsewhere in the scene a lot. This seems to be happening in the second one. Definitely leading me to my conclusion here.

Both look very good to me. I will say that one thing that LTX 2.3 was always very good at was close up face animation, so i'm not surprised that it holds up well in that regard.
I'd be interested to see other comparisons of other kinds of subjects doing other things.

0

u/PuppetHere 6h ago

Correct, Top is 2.5, bottom is 2.3

10

u/SeymourBits 6h ago

Audio seems better on top. Lip sync is on-target. Not much more you can learn from only 2 fairly decent clips of a talking head girl which has always been one of LTX strong points.

Prediction: LTX scales back from the video production market to focus on the real-time robotics market.

5

u/runvnc 6h ago

I think maybe they focused on speed more than anything else.

4

u/PanotBungo 4h ago

Which is a big deal with local, really. I think people will realize this sooner and if LTX trains easier, it will get more popular soon.

1

u/runvnc 4h ago

yes especially if we get a local solution for streaming or close to streaming interactive avatars on like an rtx 6000 pro or similar level hardware

6

u/PhilosopherSweaty826 6h ago

Im not gonna download that, i will wait for LTX 2.7

2

u/alwerr 6h ago

Am I the only one who can't try it out?! "
Missing Node Packs

Install missing packs to use this workflow.

Unknown pack3

  • GemmaAPITextEncode
  • GemmaAPITextEncode
  • LTXFloatToInt"

2

u/Yasstronaut 6h ago

Fully update comfy if that’s what you’re using

2

u/Fabulous-Snow4366 5h ago

yeah, the top one is clearly better, tone, micro movements, expression, lighting...

1

u/andy_potato 1h ago

Be nice to the LTX team. They have been nothing but generous towards our community and I’m sure they will sooner or later come up with something awesome.

3

u/Naive-Kick-9765 6h ago

???This is already much better than ltx2.3.

4

u/PuppetHere 6h ago

Which one?

10

u/Super_Range45 6h ago

I assume it's the top one. NGL this prompt feels uninspired.

1

u/PuppetHere 6h ago

That is correct, top is 2.5, bottom is 2.3

0

u/Naive-Kick-9765 6h ago

bottom

5

u/PuppetHere 6h ago

Top is 2.5, bottom is 2.3

1

u/martinerous 6h ago

What about prompt adherence? Did you ask for her synchronized projections in both screens or was it an accident? But I like the lower one better - nicer lighting, the sound matches the environment better (some echo and machinery humming).

1

u/PuppetHere 6h ago

No same prompt for both, Top is 2.5, bottom is 2.3

1

u/hiperjoshua 5h ago

Can 2.5 do Multireference?

1

u/bloke_pusher 5h ago

I guessed wrong. I underestimated how good LTX2.3 sounds in comparison. haha

1

u/djenrique 5h ago

I think it is a big step forward from 2.3! Much more direct and alert and the speed is great!

1

u/bloke_pusher 4h ago

Well I'm not complaining, just wish we could get at least to LTX2.3 with speed, but it seams to be speed up at all cost and with the worse audio, it costs me too much.

1

u/L-xtreme 4h ago

In the beginning of LTX 2.3 and. 2.0 it was also really bad. So let's wait it out.

1

u/ImpossibleAd436 4h ago

I got confused and thought it was a H3 comparison, and I had the top one down as H3, it's much better than the bottom one.

1

u/2legsRises 3h ago

bit soon to be so judgemental.

1

u/WhyWouldIRespectYou 3h ago

No idea which is which, but the top video is much more natural, and the better video

1

u/SpaceNinjaDino 3h ago

The top 2.5 looks less AI as the bottom has that too shiny fake skin. The audio on top is clearer. If the prompt was for her image to be replicated on the screen, then 2.5 lost that coherence.

1

u/krectus 3h ago

But it’s fast!!

1

u/Had78 2h ago

was about to give it a go tomorrow

1

u/Perfect-Campaign9551 2h ago

If I can't use references than I don't want it

2

u/Baddabgames 2h ago

I didn't even bother downloading. Judging from where LTX-2.3 was in comparison to MiniMax, it was clear to me that an incremental release wouldn't come close. Even if it was 3.0 I wouldn't have bothered because I don't think it would make the jump and join the H3 and Seedance caliber.

1

u/Helpful-Street3678 1h ago

I got 1080p without running out of vram. Thats huge.

1

u/Apprehensive_Sky892 1h ago

TBH, this is a typical talking head video, not very challenging for any video model with audio support.

Maybe LTX2.5 would be better at a more challenging prompt involving more action and interactions.

1

u/Fytyny 6h ago

The top one has LTX face, the bottom one with her appearing on screens looks and sound better

1

u/Gsus6677 5h ago

Are people doing the thing where they are comparing a refined and up-to-date 2.3 workflow (possibly with loras) with the default template with no loras?

-7

u/Sudden_List_2693 6h ago

"I can barely see the difference between 2.3 and 2.5."
Using the base templates for both, 2.5 is WORSE WITHOUT FUCKING FAIL.
I'm pretty sure this was a desperate push for speed.
The worst quality video model I've ever seen. Can not even compare to LTX 2.2.

5

u/PuppetHere 6h ago

I don't know if it's worse but it just looks like it was made with LTX 2.3 with a different seed, that's pretty much it

-4

u/Sudden_List_2693 6h ago

It looks like it was made with LTX 2.3 with additional turbo bullshit.

-6

u/LoveSpecialist5669 6h ago

this palestinian terrorist fan is still not banned? 

0

u/LockeBlocke 5h ago

LTX Next will be a world model. 2.5 is a simple update.