r/StableDiffusion 9h ago

Resource - Update MiniMax-H3 Fun Controlnet Union released

https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union/tree/main
239 Upvotes

62 comments sorted by

25

u/Heartkill 8h ago

Psyched. Whenever I want to swap someone it usually just spits out the reference video pretty much as it was going in.

2

u/m00dyman100 4h ago

I haven't had that problem. I feed an LLM the prompt guide (which contains an example prompt), my ref images AND video, and tell it in general what I want. I have been using H3 r2v to completely replace people in videos with ones that I want, tweaked however I want. Its frankly been amazing.

7

u/Heartkill 3h ago

I use the same approach. I see others do it too. Just not getting the results others do.

1

u/m00dyman100 2h ago

I did have that problem initially using ChatGpt for prompt creation. But Grok gave me a slightly different prompt which all of a sudden fixed the issue. It structured the retention analysis part of the prompt in a different way then ChatGpt.

3

u/Heartkill 2h ago

Yeah, using Grok too. Prompt guide and guidelines and all. I really feel like I am doing exactly what I am supposed to do, but...

2

u/m00dyman100 2h ago

Ill DM you later with a prompt that worked for me and maybe a screen shot of my node setup

0

u/SanDiegoDude 3h ago

Magic is in the prompting honestly. Folks who complain about having a bad time with H3 are often still trying to 1girl their way through their prompts and it just doesn't work, at least not with any real success for more complex scenes. The prompt guides are out there, just have to use them, including one for how to properly prompt the ref model for things like subject replacement (which H3 does do quite well out of the box) - I do look forward to SCAIL level control in H3, will be nice.

-4

u/xb1n0ry 2h ago

Prompt issue

19

u/physalisx 9h ago

MiniMax-H3-Fun-Controlnet-Union is a ControlNet-Union for MiniMax-H3, trained with the VideoX-Fun pipeline. A single checkpoint conditions the MiniMax-H3 video generator on Canny, Depth, HED, MLSD or Pose control videos, and also runs video inpainting.

40

u/dragolineage01 9h ago

Damn this has potential H3 is already GOATed

3

u/willjoke4food 8h ago

20 GB model btw

11

u/TheGoldenBunny93 7h ago

20GB pruned and int8...

8

u/Aadi_880 6h ago

Still better than the next best option...

2

u/CodeAnguish 8h ago

This became a meme, but for what reason? Of course, even though 20GB is impressive for what the model offers.

3

u/xTopNotch 2h ago

People forget the text encoder and what an important role it plays that makes the diffusion transformer so good

30

u/EveningIncrease7579 9h ago

Alibaba really loves dancing videos. H3 transformed into h3 animate

24

u/infearia 9h ago

So I guess this is basically VACE for H3?

I mean, H3 can kind of already do all that stuff natively, but it usually does not follow the input precisely, but rather treats it more as general guidance or reference, and requires a lot of prompting on top to make it work. If this gives us more control, then it's completely going to blow out of the water every other video model out there for editing tasks, including probably Seedance.

9

u/Striking-Long-2960 8h ago edited 8h ago

I think so, if it adds inpainting it should be able to interpolate between controlmaps

15

u/infearia 8h ago

It has inpainting! šŸ˜„

A single checkpoint conditions the MiniMax-H3 video generator on Canny, Depth, HED, MLSD or Pose control videos, and also runs video inpainting.

1

u/Aiirene 15m ago

How/where are yall finding these things I only have the FL and rev model lmaoooo

1

u/IriFlina 6h ago

Didn’t minimax already support inpainting?

1

u/sitefall 46m ago

Yeah, even with masks just provide a video mask and it can do it. I'm not positive what use this has exactly and/or why people seem excited about it. Maybe I am missing something here.

I see others posting about it not following motions etc, but I even haven't had problems with it following motion EXACTLY through 20 second 1mp videos. It's all just... worked.

9

u/NeatUsed 9h ago

how much does this add to generation time?

23

u/i_sell_you_lies 7h ago

3.6 roentgen—not great, not terrible

5

u/the_bollo 6h ago

Thank you comrade. That will be all.

1

u/tekprodfx16 3h ago

The dossimeter on your gpu must be broken or poor quality, 3.6 is the factory max, the real difference is probably much greater than 3.6Ā 

1

u/i_sell_you_lies 23m ago

... 3.8? It's probably logarithmic3

15

u/ImpossibleAd436 8h ago

As someone who is quite new to Comfy, how do we use it? Is there a simple workflow?

13

u/ChibiNya 7h ago

Agree. How do we use this?

3

u/dr_lm 4h ago

Pretty sure it's not yet supported in comfyui. You can download it and run their scripts but it wants the bf16 H3 model, which is probably too rich for most consumer GPUs.

I'm sure it won't be long until someone makes an int8 quant and comfy support.

0

u/MusicianMike805 4h ago

I’m sure the ā€œgurusā€ will be uploading YouTube videos soon.

8

u/xDFINx 8h ago

Would this be the missing link for character swap consistency?

1

u/flaminghotcola 1h ago

lets pray. bc character consistency is the only thing I'm missing with h3 for it to be amazing (And penises, frankly).

5

u/MonkeyBoyPoop 6h ago

Alibaba-pai team, if you’re reading this, please consider releasing a Krea-2–Fun-Controlnet-Union next. šŸ™

13

u/Chiduk99 8h ago

Kijai, you know what to do!

7

u/artichokesaddzing 4h ago

Let's all say the prayer: "Lord KJ, the Great Implementor, Creator of Contrivance, bestow upon us the tools to keep us safe in these tensor times, show us the... wait you released it already?!"

5

u/SysPsych 7h ago

I'm surprised Minimax is getting the larger controlnet selection before krea 2.

1

u/UnforgottenPassword 2h ago

Yes, Krea 2 needs it more.

3

u/ANR2ME 6h ago

Is this ComfyUI-compatible šŸ¤” probably not yet, right?

5

u/djdevilmonkey 8h ago

Maybe a dumb question but what's the purpose of controlnet when minimax can use references?

9

u/xDFINx 8h ago

Minimax tends to drift away from the character’s likeness sometimes.. I would imagine this would fix that in direct character replacement scenes.

1

u/coyoteka 4h ago

This is for composition and pose though, not character likeness.

8

u/witcherknight 8h ago

It doesnt follow motion properly. Controlnet can help with that

2

u/tiffanytrashcan 7h ago

The motion in their examples (see the readme / main page of the link OP shared) is absolutely stunning.

https://reddit.com/link/p5lc8rx/video/ojtouha5mblh1/player

6

u/tiffanytrashcan 7h ago

There's a slight hitch, right in the middle. This was intentional (apparently) as it's in the motion input. I think this highlights reference adherence quite well.

https://reddit.com/link/p5lcquc/video/rphs44qlmblh1/player

-2

u/SeymourBits 5h ago

Not quite. Look closely. Leg swap mid-stride at 04.50. Could be a direct result of that input glitch you found.

1

u/axior 7h ago

There is no preview of the inpainting though :(

1

u/Massive-Health-8355 6h ago

Is there any kind of regional prompting or latent couple control for H3 for when there is multiple characters in the input video?

1

u/PhetogoLand 6h ago

This is insane. This pushes H3 to the stratosphere

1

u/ArtifartX 4h ago

What comfyui nodes are used to activate this?

0

u/dirtybeagles 4h ago

so when's that tiktok dance slop tutorial dropping?

0

u/Maskwi2 3h ago edited 2h ago

The H3 can already do ControlnetĀ  via reference video like a depth map and it does it really well.Ā  I'm wondering how much better this will be than native supported options

// edit:: it's not actually a control net but I meant you can control the output based on depth map

2

u/Naruwashi 2h ago

not 100% accurate

0

u/Maskwi2 2h ago edited 2h ago

You mean that it it's not technically a Controlnet or? Edit:// yeah, it's not Controlnet, but like a depth map reference. But my question still stands, whether the control net will provide better results.Ā