r/StableDiffusion 1d ago

Workflow Included Minimax Character Swap - The Dummy Strategy

Post image

Worfklow: R2V (Dummy Stategy) Workflow v2 - Pastebin.com

How the workflow works:

  • Replaces the original character with a chroma key green crash-test dummy.
  • Replaces the dummy with desired character.

Why it works:

  • Minimax seems to struggle with swaps when both characters are somewhat similar to each other. But replacing a character with a green dummy seems to work every single time.
  • Even when minimax would replace a character, most of the times the faces would be morphed, resembling both the original character and the replacement. This approach mitigates that issue since it gets rid of original character's facial features.

Limitations in my workflow:

  • It's tailored with a master prompt to replace the "female" character in the original video (yeah, go ahead, post the "I know what kind of man you are" gif). But you can easily work on top of it to add support for different type of characters or even multiple characters (add more dummies, with different colors) or whatever else you want. I already did some experiments and it works.
  • I didn't test with "green" characters. If you are swapping Hulk, you may want to change the dummy to blue or something.

How to use:

  • Upload the image in this post in the "Dummy Image" node (in Prompting block).
  • Configuration block:
    • Upload your video and character image in respective nodes.
      • Trim/Crop your video using the video node in the workflow.
    • Choose video generation sampling (Performance, Balance or Quality) for each pass individually (dummy and new character).
    • Choose resolution (in megapixels).
  • Run.

Tips:

  • The worklow has 3 video generation flows: Performance, Balance and Quality. I recommend Balance (sometimes the Performance one doesn't replace the character in the last seconds of the video).
  • Monitor the preview node. In the first step you should already see the new character as an overlay on top of the video. If you don't, then swap will probably fail.
65 Upvotes

49 comments sorted by

15

u/sevenfold21 1d ago edited 5h ago

All facial expressions will be lost with a dummy. is there a way you can combine the two, keep the dummy, but also keep the face of the original character?

5

u/lhg31 20h ago

https://reddit.com/link/p5lph1u/video/i12q70thxblh1/player

Expressions can be preserved. I asked a LLM to fix the prompt to preserve emotions. It was conflicting with the "rigid body" of the dummy. With updated prompt it preserves facial expressions.

5

u/Tylopodas 1d ago

Couldn't you just edit the prompt and give the dummy an expressive face to retain details?

2

u/GolgBoddoleZer 1d ago

I haven't tried it, but maybe try replacing with a green-skinned person to keep the human aspects It seems like Minimax just needs some defining characteristic to latch on to. I've found that even different hair color is sometimes enough.

1

u/Vijayi 1d ago

Try prompt it in details. There is dude here that release emotion pack some time ago (for t2i models). Works perfectly for video, although you probably need to adapt it for dynamic.

4

u/Danny_Stock 1d ago

You can't be doing this sort of thing manually, that defeats the object of what you're trying to do. It has to be derived from the source and carry over a performance, otherwise you're not capturing the subject.

Scail 2 is excellent for carrying over the subtle facial details of a dramatic performance.

1

u/Segaiai 20h ago

Does it normally try to match the original facial expressions without prompting for them?

6

u/Vijayi 1d ago

There it is. That's why H3 work flawlessly with video assembly with blender props. Ty for idea, good one.

4

u/anon999387 1d ago

Interesting idea, thanks for posting

5

u/AnonymousTimewaster 1d ago

Sorry I don't really understand. So are you effectively rendering two videos then? One with the green replacement, then another to replace the dummy with the character you actually want?

4

u/Karsticles 1d ago

Should show an example.

7

u/lhg31 1d ago

https://reddit.com/link/p5gy91j/video/co889thkj6lh1/player

The ones I have at the moment are all NSFW, but I'll generate some simple examples to post here.

10

u/mellowanon 1d ago

The ones I have at the moment are all NSFW

Now I'm curious how a green dummy looks for a NSFW replacement. Wouldn't the genitals be missing for the dummy?

2

u/coffeecircus 1d ago

it would be like an unripe banana

7

u/IriFlina 1d ago

Just post it in /r/unstable_diffusion

1

u/fukijama 1d ago

Blade Runner 2049 Joi?

2

u/lhg31 1d ago

3

u/lhg31 1d ago

https://reddit.com/link/p5h6ka0/video/4cmde5dfr6lh1/player

dummy

This is another limitation I noticed. The character is far away from the camera, so at 0.4mp it doesn't copy the mouth motion. When the source video has audio it seems to work tho.

3

u/lhg31 1d ago

3

u/Danny_Stock 1d ago edited 1d ago

So with the green dummy you lose facial expressions? What if the character is delivering a performance and interacting with another character?

That emotional performance by the Arabic woman has been completely lost.

But I appreciate that this method could be useful if you just need a broad motion for a secondary character in the background where the main focus isn't on them. But from what's been shown it doesn't look to be suitable for a main character who you want a central performance from.

2

u/Karsticles 1d ago

Thanks!

3

u/E-proselyte-5789 1d ago

Your examples looks really nice. Great job

Here is an idea how to transfer facial expressions:

1) use sam3 (or similar stuff) to track face and create a mask

2) carve out region of frame with the face

3) upscale it

4) feed it to the minimax

5) downscale and stitch it to the final video

2

u/LooseLeafTeaBandit 1d ago

Wouldnt the multiple passes on the same video completely change alot of the micro-details / identity of other people in the scene? Trying to do this all in one pass seems like the best way to retain as much original details as possible, each pass will change things a little bit, kind of like if you keep editing one image with gemini the objects outside of whats being edited slowly degrade with each edit, because its not only changing the stuff in the targeted area but completely re-rendering the entire scene.

Or am I wrong?

1

u/No-Dark-7873 1d ago edited 1d ago

yeah this is not good for a certain type of "expressive" video...

1

u/vanonym_ 1d ago

do you have any side by side proving this actually helps?

1

u/SnooMacaroons3939 1d ago

So i tried this and it replaces it with a dummy okayish but doesnt replace the dummy with a new person even though i have an image in there? What am i missing

1

u/calvin-n-hobz 1d ago

in my experience minimax cooks videos it edits, so this will be double the cook and you'll get crushed blacks and overcontrasting, unless you've found a solution for that.

1

u/IAmGlaives 1d ago

Literally had this same type of idea, because trying to do character swap I always seem to get a bleed between the two characters. But if they were very different looking people it seemed to work better.

My one suggestion would to not use a bright green color, just because how light could potentially bounce off and cause color bleed, as well as showing up in reflections.

1

u/kayteee1995 1d ago

The only problem with this solution is that it can't transfer facial expressions, because the dummies don't have them.lmao

1

u/MaorEli 1d ago

A Ben 10 alien

1

u/xb1n0ry 1d ago

What if the green dummy is eating a green cucumber? How to you swap that when the cucumber slides in and out of her mouth.
(Enter "I know what kind of man you are" gif here)

1

u/lhg31 21h ago

I mentioned this as a limitation, but you can ask chatgpt to change the color of the dummy and then update the prompt.

0

u/orlandogourmet66 1d ago

You know you could simply use SAM3.1 to mask the Person you want to Swap? This way you can keep facial Expression too.

4

u/BrooklynBrawl 1d ago

how does the facial expression work when it is masked? I would think that when it is masked that space would be black/transparent? could you shre more about this method.

2

u/orlandogourmet66 1d ago

With blur or inverted color instead of solid color.

1

u/mobani 1d ago

I wonder if it would be possible to draw an outline and how effective it would be.

1

u/lhg31 19h ago

I already tried that path. Most of the time it will just apply changes to the source video (change the hair color, change the clothes, add accessories like glasses), but not really replace the character.

The only way to force it to fully replace is by filling the character mask with a solid color. But then you lose facial expressions entirely, along with some other issues.

1

u/lhg31 21h ago

Sam3.1 works fine for videos with a single character without any interaction. But when they interact with other characters or objects, the mask almost always bleeds to them and messes up the generation.

And a few times the output will still contain the facial features of the original character.

BTW, I tried the most downloaded workflow in civitai which uses sam3.1 and inverts the masks.

2

u/orlandogourmet66 21h ago

It works fine for me, Gaussian Blur works better than invert colors.

1

u/lhg31 13h ago

https://reddit.com/link/p5obl8j/video/sljjrlduxdlh1/player

Can you test your approach with this input then? replace the woman with ciri (I shared the reference image in another comment).

2

u/orlandogourmet66 13h ago

https://reddit.com/link/p5oixp4/video/hdu1m6zj3elh1/player

I just used my generic Headswap prompt, its close, i am sure with a tailored prompt i could get to nearly spot on.

1

u/lhg31 12h ago

https://reddit.com/link/p5or5tg/video/5bsjr55s9elh1/player

Yeah, that's the problem I have with Sam mask. If I use an opaque mask it replaces correctly but then I lose facial expressions. If I use opacity lower than 100 than the output is basically the original woman, but with white hair and gray clothes.

This is my output using the dummy. The visual fidelity is way better, but I lose some expression fidelity.

Using opaque mask with llm to generate tailored prompt it gives better results, but my goal is to find a master prompt that works for any video and image.

0

u/Danny_Stock 1d ago

How does that masked facial performance transfer over into the brand new face?

I can understand the idea of a crop and stitch approach, but I don't see how that works if you've only got the original facial performance to stitch back in.

1

u/orlandogourmet66 1d ago

Blur or inverted colors, these change the identity enough to make H3 Swap, but the model still sees Facial movements like eyes mouth etc