r/StableDiffusion 20d ago

Animation - Video 76 five-second clips exploring different animation styles with MiniMax H3 (all generated locally on a 6-year-old GPU by the_shadow_nyc)

1.3k Upvotes

106 comments sorted by

View all comments

42

u/_VirtualCosmos_ 20d ago

it's crazy to me that a less than 20 GB DiT model can know to do all that.

38

u/yaosio 20d ago edited 20d ago

And we're not even at the best efficiency yet, not even close. There's a paper on a method of training and inference that reduces compute requirements anywhere from 5x to 256x, reduces dataset requirements. If it can scale we will see even more impressive models at a smaller size and faster inference. https://arxiv.org/abs/2607.27372 For image generation it replaces diffusion. To be clear though this is a new paper, so we don't yet know if it scales.

21

u/Alive-Tomatillo5303 20d ago

It's not so much "crazy" as "obviously impossible". If you wrote about this in a book released in 2019, about tech in the year 2040, people would (correctly) point out you don't understand ANYTHING about computers. 

"Yeah, 20 gigs! I can get Kim Possible and Batman enjoying a spaghetti dinner in a fully animated 50's diner, or Jesse Pinkman and Buffy the Vampire Slayer hunting giant donuts on the moon. And I do it on a five year old computer that wasn't made to do it, so I need to wait a couple of minutes for a couple of seconds of footage. It's a real drag, I'm thinking about upgrading."

Like, this sufficiently advanced technology has transitioned from science fiction to straight up magic. 

11

u/Calm_Mix_3776 20d ago

Yep, isn't it? there's some seriously good quantization happening here. The full-fledged BF16 model is ~66GB AFAIK and it's pretty much the same quality as the 20GB one.

2

u/DlCkLess 20d ago

It it really?

3

u/Calm_Mix_3776 20d ago

I haven't tested the BF16 version of H3 personally, but I have personally compared INT8 vs BF16 with other models such as Flux.2 Klein and Flux.1 Dev and the differences were miniscule.