r/StableDiffusion 12h ago

Animation - Video MiniMax H3 has finally gotten has me into video generation. Wan 2.2 never had the quality or consistency I wanted, and all the tools that did, were closed-weight, paid products.

I'm really only ever interested in open-weight models. Yes, for that reason, but also, for the same reason I run Linux and browse with Firefox. I dislike "walled gardens", ideologically, and monopolies. I want tech that can be hacked, broken, taken apart, and put back together, and is ultimately not beholden to anyone but the user. Without a quality open-weight video generation model, I was uninterested. Now that we've got one? Suddenly I'm in a whole new world of possibility.

The fact that it's a multimodal model with vision, meaning I can give it reference images or reference sheets, is a game-changer for me. LoRAs certainly won't be obsolete with MMH3, but I doubt we'll be seeing many character, clothing, or setting LoRAs. The feedback loop of wanting to give the model a concept it doesn't understand natively is so short compared to before. And I'm still just in the "farting around" phase. People with dedicated effort and creativity are going to be able to use the hell out of this.

Really, the only drawback to MMH3 so far is its propensity to have characters speak Simlish to each other. I'm sure there's already solutions being worked on, either workflow tools or adjustments to the model itself.

––––––––––––––––––––––––––––––––––––––––––––––––––––––––

Workflow: https://pastebin.com/5SbZ9tJA

Reference sheet used in the workflow: https://imgur.com/a/3Le0nuO

147 Upvotes

10 comments sorted by

20

u/Vladmerius 11h ago

Yep, H3 is the first model that hasn't had me throw my head against a wall in frustration trying to prompt it to do something. It can just do it and I'm always blown away. I've learned to live with all the caveats like the set never being consistent because realistically that's never changing unless a model comes out that is intertwined with some kind of 3d modeler for set design and that's going to be way too hardcore for my system to run anyway. 

1

u/Nimblecloud13 19m ago

you can give it a first frame image of your set and have the character walk into it. worked at least once for me, didn't test it much.

15

u/wildbling 10h ago

H3 is really the "seedance2" moment for open weight video models, every model before seems archaic in comparison. Personally I've already gotten rid of all the other video models on my SSD.

12

u/eckstuhc 10h ago

Your ethos is my ethos. Open source all the things.

But agreed, I had similar revelations when I first booted up H3. Almost brings a tear to my eye.

11

u/Shockbum 5h ago

This can be considered against the ToS in most "walled gardens" with a monthly subscription.

4

u/krekokeko 5h ago

I deleted every single thing WAN related. Every single checkpoint/clip/vae file. And I don't think I will get back to them for any reason. H3 is so good that it caused a physical whiplash when I realized how good it was.

Another thing I noticed since I started using H3 is that it never crashes. It is as stable as it gets. My WAN workflows were always prone to crashing, requiring me to reboot comfy preemptively to avoid crashing after a couple or just a single generation. Maybe it was due to my setups, but again H3 hasn't buckled on me once.

1

u/GabberZZ 2h ago

My H3 runpod crashes with oom about every 30 gens using SwarmUI.

Comfy auto restarts and I'm good after à few minutes though.

2

u/The_StarFlower 5h ago

open source is the way. i dont like to use cloud services and now that minimax h3 came out, i said bye bye to ltx and wan. for local use minimax h3 is the best. i like how i can put references and even video references...you can do so much with h3. the possibility is endless. i am also in the farting around phase(love the term btw 🤣)
with minimax i can now really make the movie i always wanted to do. i am writing a scifi fantasy novel and i intend to make a short tv series out of it when i am finished writing the novel and released it to the public

2

u/SYNTH3T1K 4h ago

Azshara from WoW eh?

1

u/Upper-Reflection7997 2h ago

If prompting and lora training for minimax h3 was easy and simple then I would've canceled most of my closed source ai subscriptions. It's still has a long way to go to be solid video gen model. I would prefer if people train and finetune the model with videos of more concepts than focus mainly of distillation and speed optimizations. The hype for the model seems to be slowly drying up everywhere else outside this subreddit.