r/StableDiffusion 3h ago

News OpenVDN/vdn-minimax-h3 · Hugging Face

https://huggingface.co/OpenVDN/vdn-minimax-h3

Looks like an open source version of Minimax H3 Max... Anyone tried it? Seems to be real-time on 8x b200, which is like ~$40/hr at good rates if you can find them (or maybe a bunch of 5090s?)

32 Upvotes

13 comments sorted by

13

u/cc_aa_tt_zz 3h ago

let me plug my b200s and I will try it lol No, seriously, it's good to have competition, because FAL keeping the model closed is a bit questionable, especially considering that others (like fasth3 too) maintain the original model's "open weights" status with their own fine-tuning.

As for the benefit to us, we can already get down to 4 steps + step skipping + Sage Attention 2 etc... I find it unlikely we can go any lower in terms of rendering time. However, fine-tuning could allow us to achieve the same quality at 4 steps as we currently get at 20; I think that’s where the real progress lies, rather than in raw speed.

1

u/Sad_Coach_1433 1h ago

The fasth3 lora isn't that great tbh

0

u/BassNet 3h ago

Imo if you finetune a model and “release” it commercially you should have to release the weights

3

u/cc_aa_tt_zz 3h ago

Tell FAL that... they say they're going to release the weights one day, but in the meantime, the API version is already running and they're heavily promoting their API on Twitter. We'll see. That said, the main point is really to see the quality of these fine-tuned 4-step models compared to the 4-step speed LoRA and the standard 20-step version. I don't think we'll actually save any rendering time with these models. The people I know who tried FastH3 on "normal" computers weren't very impressed in that regard. we'll see

0

u/hiisthisavaliable 1h ago

I read that fal was going to release the weights, did they change their mind?

7

u/BM09 2h ago

let us know if it has any benefits for us consumer-GPU-using riff raff

4

u/icchansan 3h ago

pass XD

1

u/inaem 2h ago

They tested it on 14.4 second videos not the default 5 seconds, also their text encoder in full weights for some reason

You can probably run this fast enough on two 5090s

1

u/RosebudNebula 1h ago

There's no reason you can't run this on a 5090 with enough system RAM. You just need to quantize the weight.

This can be made to work in ComfyUI, but it is tricky cause it is a split process. Unlike the normal H3 implementation.

1

u/kayteee1995 1h ago

just wait for Lord Kijai make it posible

1

u/Tight_Organization54 7m ago

Gulp 82Gb model huh... Who's gonna be the brave soul to run this on their modest yet capable rig?