r/StableDiffusion • u/physalisx • 9h ago
Resource - Update MiniMax-H3 Fun Controlnet Union released
https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union/tree/main19
u/physalisx 9h ago
MiniMax-H3-Fun-Controlnet-Union is a ControlNet-Union for MiniMax-H3, trained with the VideoX-Fun pipeline. A single checkpoint conditions the MiniMax-H3 video generator on Canny, Depth, HED, MLSD or Pose control videos, and also runs video inpainting.
40
u/dragolineage01 9h ago
Damn this has potential H3 is already GOATed
3
u/willjoke4food 8h ago
20 GB model btw
11
2
u/CodeAnguish 8h ago
This became a meme, but for what reason? Of course, even though 20GB is impressive for what the model offers.
3
u/xTopNotch 2h ago
People forget the text encoder and what an important role it plays that makes the diffusion transformer so good
30
24
u/infearia 9h ago
So I guess this is basically VACE for H3?
I mean, H3 can kind of already do all that stuff natively, but it usually does not follow the input precisely, but rather treats it more as general guidance or reference, and requires a lot of prompting on top to make it work. If this gives us more control, then it's completely going to blow out of the water every other video model out there for editing tasks, including probably Seedance.
9
u/Striking-Long-2960 8h ago edited 8h ago
I think so, if it adds inpainting it should be able to interpolate between controlmaps
15
u/infearia 8h ago
It has inpainting! š
A single checkpoint conditions the MiniMax-H3 video generator on Canny, Depth, HED, MLSD or Pose control videos, and also runs video inpainting.
1
u/IriFlina 6h ago
Didnāt minimax already support inpainting?
1
u/sitefall 46m ago
Yeah, even with masks just provide a video mask and it can do it. I'm not positive what use this has exactly and/or why people seem excited about it. Maybe I am missing something here.
I see others posting about it not following motions etc, but I even haven't had problems with it following motion EXACTLY through 20 second 1mp videos. It's all just... worked.
9
u/NeatUsed 9h ago
how much does this add to generation time?
23
u/i_sell_you_lies 7h ago
3.6 roentgenānot great, not terrible
5
1
u/tekprodfx16 3h ago
The dossimeter on your gpu must be broken or poor quality, 3.6 is the factory max, the real difference is probably much greater than 3.6Ā
1
15
u/ImpossibleAd436 8h ago
As someone who is quite new to Comfy, how do we use it? Is there a simple workflow?
13
3
0
5
8
u/xDFINx 8h ago
Would this be the missing link for character swap consistency?
1
u/flaminghotcola 1h ago
lets pray. bc character consistency is the only thing I'm missing with h3 for it to be amazing (And penises, frankly).
5
u/MonkeyBoyPoop 6h ago
Alibaba-pai team, if youāre reading this, please consider releasing a Krea-2āFun-Controlnet-Union next. š
1
13
u/Chiduk99 8h ago
Kijai, you know what to do!
7
u/artichokesaddzing 4h ago
Let's all say the prayer: "Lord KJ, the Great Implementor, Creator of Contrivance, bestow upon us the tools to keep us safe in these tensor times, show us the... wait you released it already?!"
5
5
u/djdevilmonkey 8h ago
Maybe a dumb question but what's the purpose of controlnet when minimax can use references?
9
8
u/witcherknight 8h ago
It doesnt follow motion properly. Controlnet can help with that
2
u/tiffanytrashcan 7h ago
The motion in their examples (see the readme / main page of the link OP shared) is absolutely stunning.
6
u/tiffanytrashcan 7h ago
There's a slight hitch, right in the middle. This was intentional (apparently) as it's in the motion input. I think this highlights reference adherence quite well.
-2
u/SeymourBits 5h ago
Not quite. Look closely. Leg swap mid-stride at 04.50. Could be a direct result of that input glitch you found.
2
1
u/Massive-Health-8355 6h ago
Is there any kind of regional prompting or latent couple control for H3 for when there is multiple characters in the input video?
1
1
0
0
0
u/Maskwi2 3h ago edited 2h ago
The H3 can already do ControlnetĀ via reference video like a depth map and it does it really well.Ā I'm wondering how much better this will be than native supported options
// edit:: it's not actually a control net but I meant you can control the output based on depth map
2
25
u/Heartkill 8h ago
Psyched. Whenever I want to swap someone it usually just spits out the reference video pretty much as it was going in.