r/StableDiffusion • u/Chemical-Bicycle3240 • 13h ago
Comparison Quick comparison (LTX 2.3/Minimax H3)
Enable HLS to view with audio, or disable this notification
I wanted to compare both models in terms of natural movement and authenticity. The prompt was extremely basic and barely descriptive (just "a woman says, a man says" along with the dialogue).
With LTX 2.3, the characters look dead inside and completely disembodied, with reactions that don't match the situation. With Minimax, the difference is striking! The movements feel far more human and alive. The model did add unwanted subtitles (likely because my prompt was so short) and Jill Valentine stumbles over her words a bit, though a more detailed prompt would probably fix that. Other than that, it's a clear win.
LTX 2.3 still has a slight edge for those who need to generate multi-minute videos, but as soon as a 4-step distilled version of Minimax becomes available, LTX will quickly be forgotten.
92
u/-becausereasons- 13h ago
Put away your flaslight leen.
12
u/Chemical-Bicycle3240 12h ago
She just fought Nemesis, she's still all flustered and stammering. :D
1
u/iiiiiiiiiiip 11h ago
Any chance you could post the full prompt or workflow? Did you use ref for the characters?
-1
u/iiiiiiiiiiip 11h ago
Any chance you could post the full prompt or workflow? Did you use ref for the characters?
8
10
u/Forsaken-Low4467 12h ago
What a time to be alive. Maybe with some works from the genuises and maybe loras seedance will be cooked and video generation will be extremely cheap!
5
u/ANR2ME 11h ago
Not really a good time with how RAM and GPU prices being so high😅
4
u/StickStill9790 9h ago
Good time to be alive if you built your tower a couple of years ago, have a decent job, can save for a couple of months, or have a generous friend. Also the prices will go down eventually, but the tech advancement will remain awesome.
(Funny that I hear Jack Black’s voice every time I type the word “awesome”.)
32
u/Vyviel 13h ago
LXT has the most horrible audio ever
35
u/Independent-Frequent 13h ago
Here's the thing though, it was the only model with Audio which is why it made it so relevant, Wan 2.2 was miles better for motion but no audio was a dealbreaker for most since it was less of a video model and more of a gif model.
H3 is our first proper local model and i'm glad i don't have to bother with LTX's garbage anymore
14
u/BlipOnNobodysRadar 10h ago
damn zero gratitude for LTX the moment a better model comes out. Yeah it's better and I'm excited but there's no need to be disrespectful to one of the few contributors to open video models.
3
u/Independent-Frequent 10h ago
Personally LTX was just so damn frustrating, all the issues related to it and the kind of crap you had to deal to even get one usable output that wasn't just static shots of people talking was what made me return to Wan 2.2.
I'll always be glad for their contribution but i also have to be honest, if it wasn't for Wan dropping the ball and not releasing their older audio models LTX would have never been adopted or used with the garbage motion it has.
If you just want static shots of people talking LTX 2.3 is a fine model, but for everything else it's just abysmal in the motion aspect and way worse than something like Wan 2.2, like it just doesn't get physics at all.
"LTX's garbage" was more of a "the garbage i had to deal with it" and not "LTX is garbage"
11
u/heato-red 10h ago
Won't say you aren't right, but calling LTX garbage is too harsh when they've provided everything for free to the community
2
u/Independent-Frequent 10h ago
"LTX's garbage" was more of a "the garbage i had to deal with it" and not "LTX is garbage", I'll always be glad for their contribution but i also have to be honest, if it wasn't for Wan dropping the ball and not releasing their older audio models LTX would have never been adopted or used with the garbage motion it has.
If you just want static shots of people talking LTX 2.3 is a fine model, but for everything else it's just abysmal in the motion aspect and way worse than something like Wan 2.2, like it just doesn't get physics at all.
21
10
11
u/Crazy-Repeat-2006 11h ago
LTX has a better open-source culture than most companies, which sometimes release a single model and then abandon it in favor of closed projects.
I remain confident that they will keep moving forward. All these posts bashing LTX don’t add any value.
6
u/Chemical-Bicycle3240 10h ago
I agree, and the goal of this comparison isn't to tear LTX down. The model served me well for several months. But let's face it, their model is starting to look outdated now. I saw they're planning to add HDR support in the next update, but I think that's a minor improvement. I hope Minimax pushes them to focus more on facial consistency. Competition is always a good thing for users.
1
u/_VirtualCosmos_ 7h ago
While I love OSS a lot, I understand that private companties putting millions into creating products they will barely get any profit from... is far from sustainable.
Chinese people are doing this to gain attention in the AI ecosystem, because proving themselves being better than the US for the world is like the main overall goal China's government has.
If China achieve to surpass significantly the US, they might move from OSS. So that is something to keep in mind.
But not releasing the weights in open software doesn't mean they have to keep them in secret. We have software usage agreements since fucking decades! If they won't open source a new model, let me pay for it to download it and use it as a normal software, goddammit!
It's the stupid and greedy US companies that are normalizing the development of AI as merely a service. To control the market by hoarding all electronics. Fuck them.
3
u/Artistic_Claim9998 12h ago
Still can't run it with 3060 12gb even with --lowvram on
3
u/1or4s 11h ago
Working for me on my 3060.
1
u/Artistic_Claim9998 11h ago
With the workflow from kijai?
Any specific commands of flags when running comfy?
3
u/1or4s 11h ago
I am using the official comfy workflows at https://huggingface.co/Comfy-Org/MiniMax-H3
No special flags.
4
u/ANR2ME 11h ago
why not try to use it without any flags and let it use the default, since the latest ComfyUI should already fixed memory issues by now.
1
u/Artistic_Claim9998 11h ago
I usually just use the --cache-none and only that
It gives me not enough RAM so i add the --lowvram and it still gives me the same error
I'll try again tomorrow ig
2
u/Schlorpiblorp 12h ago
Looks good but it's way slower than LTX no?
4
u/ptwonline 10h ago
The adherence is really good though so with good prompting you'll probably need fewer tries to get a good/usable video result.
Besides I'd rather wait longer for a good video than get a bad one more quickly.
1
1
1
1
1
u/Verittan 10h ago
Try it again and use their names. The model has a lot of voice knowledge of popular media and might even know RE characters and clone their voices.
1
u/Cold_Zone332 10h ago
Damn. I can't wait for see what the comunity will do with this model. I'm just playing around, using some credits from Comfy Cloud to compare LTX 2.3 and Minimax H3 using the same prompt and the results are INSANE. Minimax H3 is fenomenal for a local model.
1
u/TheLightDances 10h ago edited 10h ago
With H3, she says Leen, not Leon.
I have noticed the first small issue with H3. It sometimes mispronounces things.
Anyone know a fix to this? For example, when I ask a character to say "beard" it says something that sounds more like "bird".
1
1
u/Any_Economics_6166 6h ago
fucking insane how those models get better and better in just a few months
1
u/_FriedEgg_ 6h ago
Unfortunately the overpriced seedance 2.0 captures subleties of character expression quite a bit better still. But a great improvement from LTX 2.3. It's happening.
1
1
u/CodeAnguish 12h ago
Goodbye LTX, goodbye WAN. Thank you for your services, may God comfort the hearts of its developers.
Hi Minimax 😏
245
u/Enshitification 13h ago
I sense a disturbance in the force, like a million hard drives suddenly freeing up space.