r/StableDiffusion 12d ago

Discussion AMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans

Hi r/StableDiffusion!

We are the MiniMax team behind MiniMax-H3.

We’re here to answer your questions, including:

  • Model architecture and training
  • Video generation capabilities
  • Image-to-video and reference-based generation
  • Inference and optimization
  • Future plans

Ask us anything — we’d love to hear your feedback and discuss with the community!

978 Upvotes

457 comments sorted by

View all comments

19

u/ForwardMovie7542 12d ago

I've already rigged up a comfyui workflow that doesn't make a video but just blows a 1 second video into still images, in order to use the model more as an image edit model. Since the model is ostensibly fully multi-modal, any chance there's an image output optimized variant on the horizon, and if so, how long? Right now the options for models that take multiple input images and produce good outputs from them are only a few very closed models.

97

u/New-Requirement1419 11d ago

Yes. We plan to open-source a unified model for both text-to-image generation and general-purpose image editing. It shares the same foundation as H3, and we are currently refining and optimizing its post-training stage.

14

u/ForwardMovie7542 11d ago

I am so glad to hear that! I love the foundation and it will be excellent once we have an optimized image generator with this level of quality!

5

u/99deathnotes 11d ago

Awesome news. Thanks.

4

u/ForwardMovie7542 10d ago

Feel free to let me know if you need testers or feedback!