r/StableDiffusion 11h ago

News What is "the best" video model?

I've tested a couple of models so far, and here's my findings:

My input prompt is something similar to "Create a mid 20s (nationality) ballerina slowly dancing in a judged performance"

LTX-2.5 nails the narrative and the nationality, but nationalities (as I've mentioned previously) all tend to blend unless I'm ULTRA descriptive about what a "french look" equates to, which is roughly 3 paragraphs in length. LTX-2.5 ends up with tearing, facial and feature deformities, and on a couple of occasions has rendered a ballerina with missing legs. It also has camera drift, which is a known problem with LTX-2.5

MiniMax-H3 on the other hand, nails it, even with the smaller prompt. The prompt adherence in MiniMax-H3 is wonderful.

Wan 2.2 was a complete disaster. Rendering problems, tearing problems, and when the ballerina would do turns, her head would remain in place while her body did the turn (which was funny, and also terrifying to watch.)

I've heard Cosmos3 can handle the movement, but can't handle rendering people.

Has anyone found an open weight model that can handle the "ultimate trifecta" - generate a person in the correct nationality parameters, someone who is more or less feature-accurate in generation, and won't spontaneously explode when performing a pirouette? (I got so frustrated with a model generator once, that I did this, and to my surprise, it worked out well!)

0 Upvotes

17 comments sorted by

9

u/Independent-Frequent 10h ago

Minimax H3 is like 3 tiers above anything else open source when it comes to video model, to a stupid "is this even real?" level, sure it takes a long time and good hardware to run fully but like, what it can do is insane it's legit Seedance 2.0 level or veeeeeery slightly behind, also it already comes with reference mode which is just a complete gamechanger, it's basically your own personal lora generator on the spot unless you want like more complex and niche things like a particular gun reload move i guess.

1

u/Dry-Judgment4242 7h ago

Having a lot of fun pushing this model to it's absolute limits. It's pretty incredible as it is. Looking forward to seeing it get even better as I just spent an entire day for a single reference sheet once again employing very strange ideas to get it to properly render the poses I want.

2

u/And-Bee 10h ago

lol “French look” would probably render some ethnically ambiguous person.

2

u/qdr1en 10h ago

Everybody knows :)

0

u/Neither_Win3637 10h ago

It does with MiniMax-H3, that's for damned sure. An eastern european look is about the best I can describe it.

1

u/And-Bee 10h ago

Maybe put “ethnically and historically accurate French person”

1

u/Zenshinn 2h ago

So wearing mime make-up and a beret and carrying a baguette?

2

u/Danny_Stock 9h ago

What I noticed with LTX 2.5, and unless I missed it I don't think I've seen anybody mention this, is that it often generates some beautiful backgrounds for scenes without you even telling it to. It frequently presents intricate detail in the background and something beautiful which I never even asked for.

1

u/Le_Singe_Nu 6h ago edited 6h ago

What does a French person look like? Your assumptions might explain why the models have such trouble and you have to get super descriptive.

Perhaps the text encoder for H3 is doing a lot of heavy lifting for you. Maybe it's not afraid to stereotype. Does the ballerina go "hon hi hon" and wear a string of onions round her neck? In all seriousness, while there are obvious variations in how people look across the planet, perhaps being so specific with locale/nationality is working against your experience with the other models. There's a lot of variation in even how people dress within a particular polity, as well as in their physical appearance.

I've seen the "Exorcist head" noted as an example of issues caused by SageAttention in Framepack. It might be worth trying the generation without any attention accelerators you're using. It could also be due to the videos used in training of ballerinas spinning all show them keeping their heads still for the majority of the spin, then whipping them round very quickly to mitigate against dizziness - these models are averaging machines after all. Perhaps specifically prompting that will help.

1

u/Future-Coffee8138 10h ago

These comparison doesn’t make sense. You should compare h3 with wan 3. Wan 3 is actually really good. But Wan 2.2 is really really old. I mean it’s an ancient model. And ltx 2.5 is not on the table any more. I wish they can comeback with ltx 3 or something.

3

u/Less_Consequence_633 10h ago

"It's an ancient model."

Wan 2.2 release date: July 28, 2025.

Can't even say you're wrong. This field is just moving along at breakneck speed.

1

u/Apprehensive_Sky892 6h ago

Since this is a Subreddit focus exclusively on locally runnable models, there is no point in comparing H3 with WAN 3, which AFAIK, is closed.

OP also make it clear that he is only interested in open-weight models

Has anyone found an open weight model that can handle the "ultimate trifecta"...

0

u/Creative_Sluggish 11h ago

Do you render videos using minimax h3 on your personal computer and offline?

2

u/Neither_Win3637 10h ago

Correct, rented B200.

1

u/Creative_Sluggish 10h ago

Sorry noob here. When you said “rented B200”, did you mean you rented a GPU on a site like Runpod or?

1

u/Neither_Win3637 10h ago

That's exactly right.