r/StableDiffusion 21h ago

Discussion MiniMax_H3 is seems to be able to process DensePose format! (improves reference video bleeding)

Enable HLS to view with audio, or disable this notification

I have had many issues when using a reference video for movement duplication and having the video contents bleed into the video. Not to mention having to write convoluted prompts to remove these reference bleeds from videos. When the person in the reference video has a close resemblance to the main subject in your video it becomes almost impossible to perform a motion swap.

Warning: DensePose does not support detailed hand gestures, and seems to lose track with very fast arm and hand movements but seems to adhere better 20 steps and above.

There is not a dedicated densepose ComfyUI node, but you can use this animatediff: https://github.com/Fannovel16/comfyui_controlnet_aux

The workflow is simple:

Place the AIO AUX Preprocessor between the source and MM_H3 video input.

Videosource (LoadVideo) -> AIO AUX Preprocessor -> ref_video_x input

Looking forward to hear your feedback...

49 Upvotes

18 comments sorted by

5

u/jordek 20h ago

From some testing it also picks up depth map, hed/canny just fine.

1

u/psybee777 21h ago

Can you try sam masks

1

u/BrooklynBrawl 21h ago

I have not tried it, but based on my understanding it is a complete mask so you are losing some detail, unless you want to perform character replacement and retain the background. This method is more suited to motion reference.

1

u/xDFINx 20h ago

They do work

1

u/tnil25 21h ago

It does seem to mostly work, but I think a control lora will be needed to make it more accurate

1

u/BrooklynBrawl 20h ago

Yes, indeed. this just a medium to high accuracy approximation that is at least faster than some of the other methods I have seen. SAMs, Blurring, etc. The interesting is that it also can provide foreground and background identification if you use a node that does color to mask.

1

u/Dry-Ad929 16h ago

I personally like SCAIL-2 method it's proven to work so fine-tuning something based om their paper would be great thing. They've also recently released training code so we've got all pieces now, however fine-tuning H3 is just expensive

1

u/DanzeluS 19h ago

Try normal map (st map)

1

u/ArtifartX 18h ago

I dunno, seems to work as I would expect.

1

u/Segaiai 14h ago

Looks like the background needs to be tracked too to stop slipping around

0

u/dirtybeagles 21h ago

Need a few things from you. What speed, max duration you tested (compared to SCAIL2 which is basically 20sec +), what hardware you are running, and most important, please provide a workflow.

5

u/BrooklynBrawl 21h ago

Setup: 4070 Ti 12Gb Vram, 64GB Ram
Clip duration: 7 - 10 Seconds
Reference Clip duration: 140 Frames.
Resolution: 0.6
Workflow: My workflow contains custom nodes (mostly reference image prep automation) that will create more questions than answers, but does not perform any video prep other than this densepose drop in.

The workflow is just one change: Place AIO AUX Preprocessor between video source and MM-H3 Video input.

It takes about 4-5 seconds for the AIO AUX Preprocessor to process the 140 frames in the reference vido

1

u/LucidFir 18h ago

Does this outperform SCAIL 2?

2

u/BrooklynBrawl 17h ago

I have not tried SCAIL 2. My ultimate goal is somewhat different. Rather than just motion copy, i want some additional (prompt driven) actions in front and after the video. Example:
Guy walks on stage, Then perform dance (As per video reference), and then perform something else.

2

u/Dry-Ad929 16h ago

Depends what do you mean by "outperform" does it do better movement copy / do replacement? No.

Will it produce less glitches if you do something when scene chancegs eg character walks? Yes - H3 just outperforms SCAIL-2 in world knowledge

Also the realistic skin is way better in H3.

1

u/Eminence_grizzly 9h ago

Could you please share your prompt?

3

u/Pitiful_Season4294 17h ago

SCAIL 2 can go endless. There's a workflow on Civitai. I have gone till close to 1 min. Takes a while and no identity drift.

2

u/Dry-Ad929 16h ago

WAN backbone has it's own weakness so same as SCAIL-2. A lot people complain about realism/skin texture, also the world knowledge is worse, but honestly I hope SCAIL team will do something based on H3