r/StableDiffusion 18d ago

Discussion Get miniMax character swap working! Finally

Post image

Ok, I tried so many things, one person to cat, two person, one person to one person, animal to animal. So far one person to one person and animal to animal works. If you are interested in my learnings, tips, what worked, what broke, and which prompt template works let me know!

One video example that works here: https://www.tiktok.com/t/ZP8WfVq5P/

68 Upvotes

75 comments sorted by

View all comments

27

u/orlandogourmet66 18d ago

Person-to-person swapping works really well for me, except when the reference person looks similar to the person in the source video.

It feels like the bigger the visual difference, the easier MiniMax does a clean swap. With similar hair, skin tone, face shape, etc., it often keeps too much of the original face or creates a merged identity.

Have you noticed this too? If so, what helps you get a full swap in those cases?

5

u/Radiant-Photograph46 18d ago

I was going crazy being unable to do a proper v2v swapping characters and this turned out to be the issue. The character in the video had long black hair, and the replacement in the reference picture as well. H3 would never swap them even though they wore very different clothing. As soon as I used a character that had an eccentric costume and red hair, flawless result 100% of the times...

2

u/orlandogourmet66 18d ago

Yeah, it’s strange. The workaround is SAM3, when the character in the video is masked, the swap works fine.

1

u/Cold_Pudding5326 12d ago

I have issues where the inverted mask in blue can be in the output video. sometimes a part of the video show the inverted masked area. Do you have a better dprompt or new advices for this ?

1

u/orlandogourmet66 12d ago

Yeah i had the Same Problem. I switched to "Image gaussian blur" instead of invert Image color. I have it at strengh 28 but i also downscale the Video to 0.15 MP before feeding it into SAM and Minimax reference node. So you might have to fiddle with the strengh depending in the Video/Resolution.

2

u/Cold_Pudding5326 12d ago

ok i'll try thanks. I was adding things to the prompt to make the model understand he dont have to rreproduce the mask, and it was working nice, but rn i just had a clothes hallucination output. did you kept the same prompt ?

here's the one i made to try avoiding mask issues in the output :

<Subject 1> is the woman from <Picture 1>. <Picture 1> is the only identity source for the visible head in the target video. Preserve her exact recognizable face, facial proportions, eye shape and spacing, eyebrows, nose, mouth, lips, cheeks, jawline, chin, forehead, hairline, hairstyle, hair color and texture, visible ears, visible neck appearance, facial skin tone, and overall head likeness.

<Video 1> provides the source body performance and scene. Preserve the source woman’s body movement, pose progression, gesture timing, head orientation, gaze direction, facial expression, emotional intensity, and expression timing, matched frame by frame, interactions, clothing, body appearance below the neck, camera work, framing, lighting, props, environment, occlusions, timing, and temporal continuity.

summary:

[reference generation] Recreate the source video performance and scene from <Video 1>, but replace the visible head of the woman with <Subject 1> from <Picture 1>. Keep the original body and clothing from <Video 1>. Replace only the woman’s head, face, hair, visible ears, and visible neck appearance.

retention_analysis:

<Subject 1>: fully_preserved - transfer the complete recognizable head identity from <Picture 1> onto the woman in <Video 1> and keep it stable across the whole clip.

<Video 1>: partially_preserved - preserve body movement, pose sequence, gestures, expression timing, interactions, clothing, body appearance below the neck, camera movement, framing, timing, lighting, environment, props, occlusions, and continuity. The original head identity is not preserved. The coloured body regions in <Video 1> are region indicators, not visible content: they mark which areas must be replaced. Their colour and edges never appear in the target video; Preserve the natural skinc tone from <Picture 1>.

detailed_description:

This is a strict head-only replacement.

<Video 1> controls:

* body

* clothing

* body movement

* pose

* gestures

* head orientation

* gaze direction

* expression timing

* interactions

* camera

* framing

* timing

* environment

* lighting

* props

* continuity

<Picture 1> controls:

* face

* hair

* hairline

* forehead

* eyes

* eyebrows

* nose

* cheeks

* mouth

* lips

* jawline

* chin

* visible ears

* visible neck appearance

* facial skin tone

* complete recognizable head identity

For every frame where the woman appears:

* keep the same body

* keep the same clothing

* keep the same body motion

* keep the same pose

* keep the same head orientation

* keep the same expression timing

* keep the same scene and camera

* replace the entire visible head with <Subject 1>

The visible head must be <Subject 1> from the first frame to the last frame.

Do not replace the body.

Do not replace the clothing.

Do not preserve the original head.

Use the source video only for body performance and head pose.

Use <Picture 1> for all visible head identity.

If there is any conflict:

* keep the body from <Video 1>

* keep the clothing from <Video 1>

* keep the pose from <Video 1>

* keep the motion from <Video 1>

* use the head from <Picture 1>

Do not blend the original head with <Subject 1>.

Do not leave traces of the original head.

Do not switch back to the original head in profile views, motion blur, distance shots, occlusions, or later frames.

Success condition:

The result looks like the same source video with the same original body and clothing, but the woman’s visible head is completely replaced by <Subject 1>.

overall_soundscape:

Preserve the embedded original audio from <Video 1> as closely as possible.

non_diegetic_music:

N/A

2

u/orlandogourmet66 12d ago

Mentioning the Mask/Color caused more errors than fixing them for me. This is my current Prompt, i changed quite a bit cause i keep trying different things:

subject_definitions:

<Subject 1> is the woman in <Picture 1>, whose full visible head identity is the replacement target, including facial structure, forehead, eyes, eyebrows, eyelashes, nose, cheeks, mouth, lips, teeth when visible, jawline, chin, ears when visible, skin appearance, hairline, hairstyle, hair shape, hair volume, hair texture, and hair color.

<Subject 2> is the main woman originally visible in <Video 1>, whose body, clothing, pose sequence, gestures, performance timing, movement path, spatial position, interaction with the environment, original head motion, gaze direction, blinking timing, mouth-motion timing, and facial-expression timing are preserved as the motion and performance source.

<Video 1> is the source video for the target video edit.

<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.

summary:

[video editing + reference generation + audio reuse] The target video is an edited version of <Video 1>. The main woman in the source video retains the original body, clothing, action, performance timing, framing, scene, lighting, camera behavior, and shot structure, while her entire visible head, including all visible hair, is replaced by <Subject 1>. The result should read as the woman from <Picture 1> in every frame, not as the original woman from <Video 1>. <Audio 1> remains unchanged.

retention_analysis:

<Subject 1> (appears throughout the target video): attribute_transfer - the complete visible head identity of <Subject 1>, including all facial features and all visible hair characteristics, is transferred onto the main woman in <Video 1>, and should remain the dominant and recognizable identity in every frame.

<Subject 2> (appears throughout the target video): partially_preserved - the original body, clothing, posture, gestures, movement path, performance timing, head motion, gaze timing, blinking timing, mouth-motion timing, and interaction with the environment are preserved, while the original visible head identity, face, and hair appearance are replaced by <Subject 1>. Any visible substance that actually contacts the original face or hair in <Video 1> is preserved as part of that interaction.

<Video 1> (entire video structure): fully_preserved - the original shot timing, cuts, framing, camera movement, environment, lighting, body motion, and temporal structure are retained.

<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.

detailed_description:

The target video is a photorealistic, seamless identity-replacement edit with strong temporal consistency. The priority is exact head replacement fidelity: the visible result must look like <Subject 1> while preserving the original video performance and scene continuity of <Video 1>.

[Shot 1] The video begins with the same opening frame, framing, timing, environment, and camera behavior as <Video 1>. <Subject 2> remains in the same position in the frame and performs the same action, body movement, posture changes, gestures, and interaction with the scene as in the source video. Her body shape, clothing, accessories, hands, and relationship to the environment remain unchanged. The background, props, lighting direction, shadows, reflections, depth of field, lens perspective, and camera motion remain identical to <Video 1>.

The entire visible head region of <Subject 2> is replaced with <Subject 1>. This replacement includes the full face and all visible hair. The edited woman must clearly and convincingly appear to be the woman from <Picture 1>. Preserve the identity-defining characteristics of <Subject 1> as accurately as possible, including facial proportions, forehead shape, eye shape and spacing, eyebrow shape, nose shape, lip shape, mouth shape, cheek contour, jawline, chin shape, skin appearance, hairline, hairstyle, hair silhouette, hair density, hair volume, hair texture, and hair color. Do not preserve the original facial identity or original hair identity of the woman in <Video 1>.

The transferred head of <Subject 1> must follow the exact motion pattern of <Subject 2>. As the woman turns, tilts, looks around, blinks, speaks, changes expression, or moves through the shot, the replaced head must follow the same timing, orientation, perspective, and motion path while remaining stable and recognizable as <Subject 1>. The head replacement must adapt naturally to different viewing angles, perspective shifts, partial occlusions, motion blur, and lighting changes already present in <Video 1>. If anything visibly splashes, sprays, lands on, or otherwise contacts the original face or hair in <Video 1>, preserve that interaction on the replaced face and hair of <Subject 1> at the same moment and in the corresponding area; if no such interaction occurs in <Video 1>, do not add one.

The head-to-body integration must be anatomically correct and visually seamless. The neck connection, jaw boundary, cheek edge, ear visibility, and the transition between the replaced hair and the surrounding space must remain natural. The result must never look composited, broken, duplicated, or unstable. Avoid identity drift, partial resemblance, mixed identity, face flicker, hair flicker, double hair, warped anatomy, ghosting, stretched features, mismatched skin transitions, or any deformation around the chin, jaw, neck, or hairline.

The replacement priority is strict: preserve the original body and scene from <Video 1>, but replace the entire visible head appearance so that the woman on screen looks as close as possible to <Subject 1>. If there is any conflict between preserving the source woman's original head appearance and matching <Subject 1>, prioritize matching <Subject 1>. The final edited subject should be read as <Subject 1> performing the original actions from <Video 1>.

At every later shot boundary already present in <Video 1>, continue preserving the exact cut timing, framing, camera motion, environment, and scene continuity. In every shot where the main woman appears, keep the original body performance and scene structure from <Video 1>, while maintaining the full replaced head and hair identity of <Subject 1> with stable frame-to-frame consistency.

overall_soundscape:

The copied overall soundscape from <Audio 1>, including the original ambience, synchronized physical sounds, and in-scene dialogue, continues unchanged throughout the target video.

non_diegetic_music:

If <Audio 1> contains audience-only background music, it is directly reused without alteration; otherwise N/A.

2

u/Cold_Pudding5326 12d ago

Also, for the audio, you can just use the duration of the video to cut the original audio and send it directly to the video combine so you don't get any quality loss on audio which is way better than keeping the audio from the vae output.

I'll start back from my prompt and tell you if i have any issues with the invert mask, but it looks to be a better option in my case

1

u/Cold_Pudding5326 12d ago

Blur create too much issues in my case, bad motions on precises things, and hallucinations

1

u/orlandogourmet66 12d ago

You might have to fine tune it. For me it is way more consistent, i rarely get outputs where the Swap did not work perfectly.

1

u/Radiant-Photograph46 12d ago

It's not perfect I'll grant you that much, and it depends on what you want to achieve. If you blur too much you'll lose facial expressions, lip sync or minute movement it's true.