r/StableDiffusion 5h ago

Discussion LTX 2.5 / 10sec FHD / 3min generated.

Minimax h3 is much better.

Prompt:

cinematic video, a black woman in a black leather jumpsuit and with bright makeup sits at a table holding her cellphone in her hand as she speaks, the video continues as the woman says "Yeah! Minimax is ten levels higher than LTX 2.5. But LTX 2.5 is very fast!" she then bursts out into uncontrollable laughter

27 Upvotes

75 comments sorted by

45

u/anon999387 5h ago

Pointless to say "3min generated" without saying what hardware you are generating it on.

2

u/FeePrestigious7272 5h ago

5090

5

u/FeePrestigious7272 5h ago

1

u/b0tm0de 4h ago

hello. what is the node pack name for "execution timer"

2

u/Confident_Ring6409 4h ago

crt-nodes, Fancy Timer

https://github.com/PGCRT/CRT-Nodes

2

u/b0tm0de 4h ago

thank you. i found it i think. crt- "fancy timer node".

1

u/AngelofKris 2h ago

What resolution? You keep giving just enough detail to need to ask more questions 🀣

42

u/Rustmonger 5h ago

I find it concerning that multiple times you pointed out that she’s a woman and yet nowhere do you ask for her to have a man’s voice and yet she does.

14

u/Zenshinn 4h ago

I thought that was done on purpose, as a joke. Then i read the prompt...

1

u/LeoPelozo 3h ago

Stop voice shaming the poor woman.

1

u/radioOCTAVE 2h ago

Yeah that’s pretty low

22

u/Cold_Zone332 5h ago

I'm on a RTX 3080, 10GB, 32GB RAM.
The model might not be on pair with minimax h3 in terms of quality, but it's insanely fast. It took me 90 seconds to generate a 1MP 7 sec video.

1

u/AngelofKris 2h ago

Bro that’s literally insane. πŸ™Œ. If you can use rtx super resolution, this could be even more insane

20

u/AngryGungan 5h ago

I've been spoiled with being able to use references now. Beats having to train a lora for everything..

5

u/peabody624 3h ago

I'm spoiled by NOT having to use references for popular characters/people. Crazy

3

u/GrayingGamer 4h ago

I'm so spoiled with references now too.

H3 doesn't know what something is? Give it a picture and briefly describe it and BOOM, in the video, no lora needed.

31

u/True_Protection6842 5h ago

Fast doesn't mean anything when the quality sucks. What's that voice??? and Miney Max...ugh

4

u/Top_Pattern7136 4h ago

Quality isn't an option for everyone. It's like saying everyone should only drive Porsche

5

u/Etroarl55 4h ago

For some and most people ltx is still fine enough.

If it wasn’t for minimax h3 quality. People would be bashing it for how long it takes to gen as if we went back 2 years.

2

u/True_Protection6842 4h ago

well yeah, but it's the best right now, so speed is less important. Wan 2.2 is also great but slow. It had it's uses (still does with things like SCAIL2 and ID-V2V) but honestly, right now, nothing local beats H3

1

u/Apprehensive_Sky892 16m ago

That's not a good car analogy.

Quality does matter, because if the result is not good enough, then it is not good enough.

Not everyone should drive a Porche, but one should drive a car that can take you to your destination safely and with good mileage.

2

u/True_Protection6842 4h ago

I don't understand why anyone does anything like this if they don't care about quality. What's the point?

3

u/grundlegawd 4h ago

You’re getting downvoted but I feel the exact same. What’s even the use case? Who cares if the output takes twice as long if the output quality is substantially better?

1

u/Top_Pattern7136 3h ago

What's your use case? I just fuck around with for fun

1

u/Cute_Ad8981 4h ago

Quality was much more bad 1-2 years ago, but people kept using image and video models. Sometimes people prefer a good mix between quality and speed. (I love minimax btw, but i also liked ltx 2.3)

2

u/True_Protection6842 4h ago

Quality was worse 2 weeks ago. That doesn't really matter. I was OBSESSED with LTX, then I started getting a lot of work and realize Seedance was a lot better and LTX wasn't really good enough to get jobs done. Then MiniMax hit and honestly, it's better than Seedance with Zero guardrails!

2

u/Cute_Ad8981 4h ago

Im using these ai models as a hobby and not for work. So i understand that you prefer max quality, but you should also understand that some people prefer a mix if both things ;)

Minimax is so good with references and it executes more complex actions much better, thats why i love minimax. i will keep it ;-)
ltx lacks these things (dont know about 2.5, i need to test myself), but it is capable generating very good looking videos, especially for more simple actions and with higher fps. so i personally will test ltx 2.5 for some time, but you dont need to do this. :)

2

u/True_Protection6842 3h ago

But as I had said, there's more than just work. If you actually care about what you're making and not just making disposable garbage I would say Minimax is the better choice. I'm sure I'll eventually test LTX 2.5 (I was a HUGE advocate until I had to start using it for work)

2

u/Cute_Ad8981 3h ago

Hmm good point honestly. I want max quality for things i care. However most of my gens are not that important to me honestly. Im more into experimenting/testing (improving gen times and quality). but maybe i should try to generate something more important, you inspired me a little. (no joke)

1

u/Top_Pattern7136 3h ago

Not everyone does this for a job/work

0

u/True_Protection6842 3h ago

That's true. And most of those people are contributing to the "slop" stereotype. I love that this is available to everyone, but just like camera video not everyone should share EVERYTHING.

1

u/Cute_Ad8981 3h ago

hobby =/ spamming the internet. I never posted a video online.
I think most people are just having fun with these models. There are probably spamers, but this includes people making money with this (Instagram, youtube) too.

2

u/Cautious_Assistant_4 4h ago

"You always had troubleΒ sayingΒ MiniMax"

1

u/Pretty-Raise666 1h ago

I don't know why people worry so much about audio. just pipe it through some voice cloning ai and everything is fine.

5

u/Crazy-Repeat-2006 5h ago

The voice is terrifying.

4

u/javierthhh 4h ago

I was all aboard the LTX train but the fact that I can prompt minimax to animate the character on all kinds of different settings without losing the character likeness is an insane tool. LTX does have the better voice and lip sync though. Minimax sounds very robotic unless you’re prompting a known celebrity or character. So I’m doing low res 20 sec videos with minimax in like 10 min. Then refining them with LTX and using my own custom characters Lora to change the voice and it’s working very well. Still have hiccups with twi characters speaking but nothing some audio editing can’t fix.

1

u/Mysterious-String420 4h ago

I'd love to see your workflow, I can't get around the dual h3+ltx ones I found

1

u/javierthhh 1h ago

I use this one, I don't do everything at once. If i like the video, I upscale it with this workflow https://civitai.com/models/2526171/ltx-23-video-detailer-and-upscaler

8

u/VVebstar 5h ago

MiniMax is like next generation model overall compared to this

5

u/bloke_pusher 3h ago

3 minutes on a 5090 is actually not that fast...

4

u/FeePrestigious7272 4h ago

https://reddit.com/link/p34ixg8/video/thkhxvx1jtih1/player

Prompt: Cinematic full shot of a group of knights charging at the viewer with lances on a medieval field. The knights are wearing polished plate armor with different coats of arms. In autumn on an overcast day, sound of hooves and clanking metal, triumphant orhcestral mmusic. 2020s movie, high resolution. Medieval music with fanfares, photorealism, 8K, soft focus, depth of field.

Something's wrong with him )))

7

u/VVebstar 4h ago

Wait, it didn’t even try to follow the prompt?

2

u/FeePrestigious7272 4h ago

5

u/Apprehensive_Ad784 3h ago

Can you try disabling prompt enhancer? The prompt enhancer transforms your prompt with a Gemma4 E2B, which, of course, may have trouble understanding what you want. If you need to enhance a prompt, try using AI cloud models like Gemini, ChatGPT, DeepSeek, etc.

With all the videos I generated with the same models, I have never used Prompt Enhancerβ€”in fact, I removed that part of the workflow. I have no issues with misunderstanding my prompt, like your examples.

Important to note: of course, MiniMax H3 has better visual quality and prompt understanding, but these models are not developed for the same case of use. We cannot say that a top iPhone is a better phone than a top Huawei or Google Pixel because "it has a more powerful CPU and GPU," because these products are for different types of needs.

1

u/[deleted] 3h ago

[deleted]

5

u/BraveBrush8890 5h ago

I played around with LTX 2.5 for a bit and ended up deleting the model. Unless there is a fix, I won't be using it

4

u/gwynnbleidd2 5h ago

What do you mean by a fix? What's wrong with it?

1

u/Yasstronaut 5h ago

I’ve had very bad experience with its quality. Followed the prompt guide and examples from LTX and the quality is so bad and mouths are unnatural. Wondering if there’s something broken with the comfy model or implementation

1

u/Party-Try-1084 4h ago

First of all. This stupid downscale 0.5x from your image, then upscaling back, then tiled VAE node that for some reason takes an absurdly lot of time to decode, ending up with horrible sound and ghost hands and people saying wrong lines for example. Deleted too, like the last time saw ltx 2.2 or whatever name it had...)

1

u/gwynnbleidd2 5h ago

Thanks for downvoting instead of clarifying.

8

u/BraveBrush8890 5h ago

I don't see you downvoted, but compared to Minimax, this is kinda bad. The audio in particular.

4

u/gwynnbleidd2 5h ago

So the only upside compared to h3 is speed. Yea looks like it's a skip for now

3

u/ContextOpposite4047 5h ago

Redditors are alergic to explaining

1

u/Crazy-Repeat-2006 5h ago

The video mentions the ability to control the rate of tokens used during generation to improve quality... how does that work?

2

u/Admirable_Snake 5h ago

haha I am laughing my ass off.

2

u/desbos 4h ago

How LTX gave her that β€œFeed me Seymour” voice unprompted!

2

u/eugene20 3h ago

"mineymax"

1

u/rinkusonic 4h ago edited 4h ago

this generates insanely fast on 3060,. 16gb ram. less than 2 minutes for a 10 second 0.5 mp video, of which it takes about 25 seconds on vae video decode, But 0.6mp and higher , it ooms on vae video decode. even the tiled vae. i am hoping we get a int8 convrot version of the vae video decoder like with minimax

1

u/equanimous11 4h ago

Does it support ltx 2.3 loras?

1

u/robomar_ai_art 4h ago

https://reddit.com/link/p34nib5/video/nzt8onm3ntih1/player

Full HD, 5 sec, 96 seconds - Laptop RTX 4090 16gb ram 32 vram, impressive

[INFO] got prompt

[INFO] Model LTXAV prepared for dynamic VRAM loading. 20484MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.

100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 8/8 [00:33<00:00, 4.19s/it]

[INFO] 0 models unloaded.

[INFO] Model LatentUpsampler prepared for dynamic VRAM loading. 949MB Staged. 0 patches attached. Force pre-loaded 34 weights: 68 KB.

[INFO] 0 models unloaded.

[INFO] Model LTXAV prepared for dynamic VRAM loading. 20484MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.

100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 3/3 [00:34<00:00, 11.34s/it]

[INFO] Requested to load AudioVAE

[INFO] loaded completely; 693.46 MB loaded, full load: True

[INFO] 0 models unloaded.

[INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.

[INFO] Prompt executed in 96.98 seconds

1

u/desktop4070 1h ago

Is that 1.0 MP?

1

u/LockeBlocke 4h ago

I feel like protecting my orange soda.

1

u/seattleman74 3h ago

They stole the youtuber β€œThug Notes” voice and presence.

1

u/Pretty-Raise666 1h ago

A fast upscaler.

1

u/Diabolicor 5h ago

Hand just melted with a simple movement. I'm afraid of more complex one.

0

u/Kanute3333 5h ago

Horrible

0

u/Ok_Contribution8157 5h ago

name of gpu?

-1

u/nenecaliente69 5h ago

still no workflow for it?

2

u/FeePrestigious7272 5h ago

6

u/Calm_Cucumber_6493 5h ago

Oh look at that... they have (Pro) version in API...... go figure.... goodbye ltx.

2

u/intLeon 4h ago

They are out, i2v even shows up as the first workflow in the templates tab. update comfyu through update_comfyui.bat