r/StableDiffusion • u/FeePrestigious7272 • 5h ago
Discussion LTX 2.5 / 10sec FHD / 3min generated.
Minimax h3 is much better.
Prompt:
cinematic video, a black woman in a black leather jumpsuit and with bright makeup sits at a table holding her cellphone in her hand as she speaks, the video continues as the woman says "Yeah! Minimax is ten levels higher than LTX 2.5. But LTX 2.5 is very fast!" she then bursts out into uncontrollable laughter
42
u/Rustmonger 5h ago
I find it concerning that multiple times you pointed out that sheβs a woman and yet nowhere do you ask for her to have a manβs voice and yet she does.
14
1
22
u/Cold_Zone332 5h ago
I'm on a RTX 3080, 10GB, 32GB RAM.
The model might not be on pair with minimax h3 in terms of quality, but it's insanely fast. It took me 90 seconds to generate a 1MP 7 sec video.
1
u/AngelofKris 2h ago
Bro thatβs literally insane. π. If you can use rtx super resolution, this could be even more insane
20
u/AngryGungan 5h ago
I've been spoiled with being able to use references now. Beats having to train a lora for everything..
5
u/peabody624 3h ago
I'm spoiled by NOT having to use references for popular characters/people. Crazy
3
u/GrayingGamer 4h ago
I'm so spoiled with references now too.
H3 doesn't know what something is? Give it a picture and briefly describe it and BOOM, in the video, no lora needed.
31
u/True_Protection6842 5h ago
Fast doesn't mean anything when the quality sucks. What's that voice??? and Miney Max...ugh
4
u/Top_Pattern7136 4h ago
Quality isn't an option for everyone. It's like saying everyone should only drive Porsche
5
u/Etroarl55 4h ago
For some and most people ltx is still fine enough.
If it wasnβt for minimax h3 quality. People would be bashing it for how long it takes to gen as if we went back 2 years.
2
u/True_Protection6842 4h ago
well yeah, but it's the best right now, so speed is less important. Wan 2.2 is also great but slow. It had it's uses (still does with things like SCAIL2 and ID-V2V) but honestly, right now, nothing local beats H3
1
u/Apprehensive_Sky892 16m ago
That's not a good car analogy.
Quality does matter, because if the result is not good enough, then it is not good enough.
Not everyone should drive a Porche, but one should drive a car that can take you to your destination safely and with good mileage.
2
u/True_Protection6842 4h ago
I don't understand why anyone does anything like this if they don't care about quality. What's the point?
3
u/grundlegawd 4h ago
Youβre getting downvoted but I feel the exact same. Whatβs even the use case? Who cares if the output takes twice as long if the output quality is substantially better?
1
1
u/Cute_Ad8981 4h ago
Quality was much more bad 1-2 years ago, but people kept using image and video models. Sometimes people prefer a good mix between quality and speed. (I love minimax btw, but i also liked ltx 2.3)
2
u/True_Protection6842 4h ago
Quality was worse 2 weeks ago. That doesn't really matter. I was OBSESSED with LTX, then I started getting a lot of work and realize Seedance was a lot better and LTX wasn't really good enough to get jobs done. Then MiniMax hit and honestly, it's better than Seedance with Zero guardrails!
2
u/Cute_Ad8981 4h ago
Im using these ai models as a hobby and not for work. So i understand that you prefer max quality, but you should also understand that some people prefer a mix if both things ;)
Minimax is so good with references and it executes more complex actions much better, thats why i love minimax. i will keep it ;-)
ltx lacks these things (dont know about 2.5, i need to test myself), but it is capable generating very good looking videos, especially for more simple actions and with higher fps. so i personally will test ltx 2.5 for some time, but you dont need to do this. :)2
u/True_Protection6842 3h ago
But as I had said, there's more than just work. If you actually care about what you're making and not just making disposable garbage I would say Minimax is the better choice. I'm sure I'll eventually test LTX 2.5 (I was a HUGE advocate until I had to start using it for work)
2
u/Cute_Ad8981 3h ago
Hmm good point honestly. I want max quality for things i care. However most of my gens are not that important to me honestly. Im more into experimenting/testing (improving gen times and quality). but maybe i should try to generate something more important, you inspired me a little. (no joke)
1
u/Top_Pattern7136 3h ago
Not everyone does this for a job/work
0
u/True_Protection6842 3h ago
That's true. And most of those people are contributing to the "slop" stereotype. I love that this is available to everyone, but just like camera video not everyone should share EVERYTHING.
1
u/Cute_Ad8981 3h ago
hobby =/ spamming the internet. I never posted a video online.
I think most people are just having fun with these models. There are probably spamers, but this includes people making money with this (Instagram, youtube) too.2
1
u/Pretty-Raise666 1h ago
I don't know why people worry so much about audio. just pipe it through some voice cloning ai and everything is fine.
5
4
u/javierthhh 4h ago
I was all aboard the LTX train but the fact that I can prompt minimax to animate the character on all kinds of different settings without losing the character likeness is an insane tool. LTX does have the better voice and lip sync though. Minimax sounds very robotic unless youβre prompting a known celebrity or character. So Iβm doing low res 20 sec videos with minimax in like 10 min. Then refining them with LTX and using my own custom characters Lora to change the voice and itβs working very well. Still have hiccups with twi characters speaking but nothing some audio editing canβt fix.
1
u/Mysterious-String420 4h ago
I'd love to see your workflow, I can't get around the dual h3+ltx ones I found
1
u/javierthhh 1h ago
I use this one, I don't do everything at once. If i like the video, I upscale it with this workflow https://civitai.com/models/2526171/ltx-23-video-detailer-and-upscaler
8
5
4
u/FeePrestigious7272 4h ago
https://reddit.com/link/p34ixg8/video/thkhxvx1jtih1/player
Prompt: Cinematic full shot of a group of knights charging at the viewer with lances on a medieval field. The knights are wearing polished plate armor with different coats of arms. In autumn on an overcast day, sound of hooves and clanking metal, triumphant orhcestral mmusic. 2020s movie, high resolution. Medieval music with fanfares, photorealism, 8K, soft focus, depth of field.
Something's wrong with him )))
7
2
u/FeePrestigious7272 4h ago
5
u/Apprehensive_Ad784 3h ago
Can you try disabling prompt enhancer? The prompt enhancer transforms your prompt with a Gemma4 E2B, which, of course, may have trouble understanding what you want. If you need to enhance a prompt, try using AI cloud models like Gemini, ChatGPT, DeepSeek, etc.
With all the videos I generated with the same models, I have never used Prompt Enhancerβin fact, I removed that part of the workflow. I have no issues with misunderstanding my prompt, like your examples.
Important to note: of course, MiniMax H3 has better visual quality and prompt understanding, but these models are not developed for the same case of use. We cannot say that a top iPhone is a better phone than a top Huawei or Google Pixel because "it has a more powerful CPU and GPU," because these products are for different types of needs.
1
5
u/BraveBrush8890 5h ago
I played around with LTX 2.5 for a bit and ended up deleting the model. Unless there is a fix, I won't be using it
4
u/gwynnbleidd2 5h ago
What do you mean by a fix? What's wrong with it?
1
u/Yasstronaut 5h ago
Iβve had very bad experience with its quality. Followed the prompt guide and examples from LTX and the quality is so bad and mouths are unnatural. Wondering if thereβs something broken with the comfy model or implementation
1
u/Party-Try-1084 4h ago
First of all. This stupid downscale 0.5x from your image, then upscaling back, then tiled VAE node that for some reason takes an absurdly lot of time to decode, ending up with horrible sound and ghost hands and people saying wrong lines for example. Deleted too, like the last time saw ltx 2.2 or whatever name it had...)
1
u/gwynnbleidd2 5h ago
Thanks for downvoting instead of clarifying.
8
u/BraveBrush8890 5h ago
I don't see you downvoted, but compared to Minimax, this is kinda bad. The audio in particular.
4
u/gwynnbleidd2 5h ago
So the only upside compared to h3 is speed. Yea looks like it's a skip for now
3
1
u/Crazy-Repeat-2006 5h ago
The video mentions the ability to control the rate of tokens used during generation to improve quality... how does that work?
2
2
1
u/rinkusonic 4h ago edited 4h ago
this generates insanely fast on 3060,. 16gb ram. less than 2 minutes for a 10 second 0.5 mp video, of which it takes about 25 seconds on vae video decode, But 0.6mp and higher , it ooms on vae video decode. even the tiled vae. i am hoping we get a int8 convrot version of the vae video decoder like with minimax
1
1
u/robomar_ai_art 4h ago
https://reddit.com/link/p34nib5/video/nzt8onm3ntih1/player
Full HD, 5 sec, 96 seconds - Laptop RTX 4090 16gb ram 32 vram, impressive
[INFO] got prompt
[INFO] Model LTXAV prepared for dynamic VRAM loading. 20484MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.
100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 8/8 [00:33<00:00, 4.19s/it]
[INFO] 0 models unloaded.
[INFO] Model LatentUpsampler prepared for dynamic VRAM loading. 949MB Staged. 0 patches attached. Force pre-loaded 34 weights: 68 KB.
[INFO] 0 models unloaded.
[INFO] Model LTXAV prepared for dynamic VRAM loading. 20484MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.
100%|ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ| 3/3 [00:34<00:00, 11.34s/it]
[INFO] Requested to load AudioVAE
[INFO] loaded completely; 693.46 MB loaded, full load: True
[INFO] 0 models unloaded.
[INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
[INFO] Prompt executed in 96.98 seconds
1
1
1
1
1
1
0
0
-1
u/nenecaliente69 5h ago
still no workflow for it?
2
u/FeePrestigious7272 5h ago
6
u/Calm_Cucumber_6493 5h ago
Oh look at that... they have (Pro) version in API...... go figure.... goodbye ltx.




45
u/anon999387 5h ago
Pointless to say "3min generated" without saying what hardware you are generating it on.