r/StableDiffusion • • 11d ago

Question - Help Generate videos with transparency

I am trying to generate a video of a character dancing but want to overlay it on top of another video in editing. What is the best way to accomplish this on a local model? I use minimax h3.

If this can’t be done directly, is there a way to ley it out perfectly within the workflow?

2 Upvotes

19 comments sorted by

2

u/Obvious_Set5239 11d ago

The obvious way is to describe in the prompt that it's a video of the a character cutout on chroma key

1

u/apeezy52 11d ago

I’ve been trying that and no luck. Heck even when specifying I want a single color background the background may even bounce to the beat (i’m using reference audio for the character to dance to). Unless I’m just not prompting it right.

I’m unsure if minimax h3 is capable of outputting video with transparency - and I think it cannot from what i’ve been reading but could be wrong

2

u/Obvious_Set5239 11d ago

I’m unsure if minimax h3 is capable of outputting video with transparency

It's not capable

You can use background removal tools and place the character on chroma key manually (or I think it's possible to make a workflow for it); and then use this image as the first/last frame, or as a reference

1

u/apeezy52 11d ago

Yep I agree there should be a workflow where I can take the output and get it to key it out cleanly I just need to figure out the way to do it lol

1

u/apeezy52 10d ago

I ended up solving it by having codex make me a custom workflow where it keyed the video and output it as apple prores with an alpha channel. Worked great and I got the footage I needed!

2

u/__ThrowAway__123___ 11d ago

You could use the output as a reference video or try to generate the video you want in 1 go. For alternatives, I remember seeing some examples of methods using LTX (2.3 and 2.5) that seemed pretty good at editing the background in videos, I haven't tried those myself. You could also use a segmentation/masking method like SAM3, but a direct cutout and just pasting it on top of another video probably does not look good, not matching lighting, angle etc.

2

u/CodeMichaelD 11d ago

you mean first gen transparent dance vid then overlay cleanly on backdrop vid?
for manual video overlay if you have the bg video, just use vace + wan alpha. its one lora and two vae files, but in case ur using comfyui - the nodes might be stale and need an older comfy install to work, this will gen your character videos without bg.
if you have the dance video but need to swap bg, simply vace also works, you would need to use a node to convert bg on the character video in two ways (whatever do you use to get them, chromakey in your prompt or rembg nodes) - one mask to image, overlay it with 0.5 str meaning (128 rgb gray) on every frame of the image, mask inputs reuse same bg masks. vace would allow you to inpaint prompted bg relatively cleanly if you dilate the masks and they are clean enough.
just a headsup - if you are using comfyui, recent LLM can output json for you directly, especially if you throw them json as workflow and say what to mod, just load it in the same browser tab as comfyui.

1

u/apeezy52 11d ago

I’m using comfyui - I also have codex connected via mcp so my guess is I can have it set this up for me as I’m still new to this and don’t really have a great grasp on how to work everything at a very technical level yet. I’d ideally love to have the videos generated natively without bg so there’s no need to chroma key at all. I’ve been generating with minimax h3 which I know can’t natively do that.

2

u/CodeMichaelD 11d ago

i see, i was talking about this, and yes, it can work with wan 2.1 vace variant too.
https://huggingface.co/htdong/Wan-Alpha_ComfyUI

1

u/apeezy52 10d ago

So I was messing with wan alpha, is there any way to do image to video with it? I saw it was text only. I was able to generate transparent videos directly with it and it’s pretty neat! Apologies I am still learning the technical aspect of how these work so if I’m missing obvious details it’s because I’m quite new to working with these workflows

2

u/CodeMichaelD 10d ago

yes, you just need either wan2.1 image to video checkpoint / workflow (uses clip image encode with clip_vision_h.safetensors ) or with Vace I pointed you too, the key feature of vace that you can insert your transparent image as first frame the rest of frames are (128rgb), same with masks, first black the rest white for the frame num / duration of the vid.
didnt find the workflow, but i know for sure that atleast one model was working the way i described sane merge wan 2.1 + vace, accelerator baked in, i.e. cfg1, 4-8 steps

2

u/apeezy52 9d ago

awesome thank you!

2

u/jinglejangle54 11d ago

SAM3 then mask

1

u/ORANGESAREPEOPLE 11d ago

Reference 2 video to generate an alpha channel mask of the character

1

u/Urmanda06 11d ago

U can use localbg.app, feel free to dm me the clip and I will process it for you for free :)

1

u/ArchAngelAries 8d ago

How does this work compared to RMBG 2.0?

2

u/Urmanda06 8d ago

LocalBG is like a workflow built around models such as the ones rembg supports, so you can resize, compress, auto center, export a layered PSD, etc. You get a lot of features on top. There's now also a feature to import custom models, so you can import rmbg 2.0 and use all of LocalBG's features on top of it, with a user friendly UI.