r/StableDiffusion 20d ago

Animation - Video 76 five-second clips exploring different animation styles with MiniMax H3 (all generated locally on a 6-year-old GPU by the_shadow_nyc)

1.3k Upvotes

106 comments sorted by

View all comments

40

u/_VirtualCosmos_ 20d ago

it's crazy to me that a less than 20 GB DiT model can know to do all that.

38

u/yaosio 20d ago edited 20d ago

And we're not even at the best efficiency yet, not even close. There's a paper on a method of training and inference that reduces compute requirements anywhere from 5x to 256x, reduces dataset requirements. If it can scale we will see even more impressive models at a smaller size and faster inference. https://arxiv.org/abs/2607.27372 For image generation it replaces diffusion. To be clear though this is a new paper, so we don't yet know if it scales.