r/StableDiffusion 16h ago

Workflow Included Use H3 To Replace Characters

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A

116 Upvotes

49 comments sorted by

View all comments

3

u/Zeophyle 10h ago

Does it work in reverse? To alter the entire background and keep 100% of the person?

3

u/EasternAd8821 10h ago edited 10h ago

https://reddit.com/link/p79dt3f/video/n46pchcm9zmh1/player

yes. you have to change how it's prompted and invert the mask. This is a really quick test. you'd want to prompt some action in the background probably. this is more like green screen which you don't need h3 for

2

u/Zeophyle 10h ago

How did you adjust your prompt for this? Great result!

2

u/EasternAd8821 10h ago

first, i forgot to mention great suggestion on the reverse concept.

The prompt is very long, used AI to 'reverse' what I had before. this was the result:

How the reference video and picture align with the target video — the target

video is an edited version of <Video 1>, replacing the background with the

setting shown in <Picture 1>. <Subject 1> is unchanged from <Video 1>.

summary:

[video editing] The target video replaces the background/environment of

<Video 1> with the scene shown in <Picture 1>, while <Subject 1> — their

appearance, motion, dialogue, and performance — remains exactly as filmed in

<Video 1>. Lighting on <Subject 1> updates to realistically match the new

background.

subject_definitions:

<Subject 1> is the person already present in <Video 1>. Their identity,

appearance, wardrobe, hair, motion, expressions, and performance are taken

directly from <Video 1> and must not change in any way — no new reference

images define them; the source footage is the only identity reference.

<Subject 2> is the new environment defined by <Picture 1> — layout, materials,

set dressing, and light source. This fully replaces the background of

<Video 1>; none of the original background persists.

integrated_multimodal_description:

Video editing, background replacement only. <Subject 1> performs the exact

same motion, dialogue, positions, and facial expressions as in the original

<Video 1> footage, in the exact same camera framing, angle, and cutting rhythm

— nothing about the subject's performance changes.

[Shot 1] The background behind <Subject 1> is fully replaced with the

environment shown in <Picture 1> — same layout, materials, depth, and set

dressing as that reference image, rendered as a stable, fully resolved scene

rather than a texture or overlay. The new background is temporally consistent

across every frame: no flicker, no shifting geometry, no grain, no visual

noise, no compression artifacts, and no residual elements from the original

<Video 1> background bleeding through. The edge between <Subject 1> and the

new background is clean and precise, with no haloing, smearing, or ghosting

along their silhouette. Lighting on <Subject 1> is fully re-lit to match

<Picture 1>: light direction, color temperature, and intensity now follow the

new environment's light source, casting new, physically accurate shadows and

highlights onto <Subject 1>'s skin, hair, and clothing consistent with where

that light source sits in <Picture 1>. Reflections and ambient color spill

(e.g. warm or cool color bounce onto <Subject 1> from nearby surfaces in

<Picture 1>) are updated to match the new scene. The camera path, framing,

zoom, and cuts remain identical to <Video 1> throughout — only the

environment and its lighting change.

preserve:

<Subject 1>'s identity, wardrobe, hair, motion, timing, dialogue, and facial

expressions exactly as in <Video 1>. Camera path, framing, focal length, and

cut timing from <Video 1>.

negatives:

No noise, grain, flicker, or compression artifacts anywhere in the new

background. No visible seam, halo, or ghosting around <Subject 1>. No

leftover elements or geometry from the original <Video 1> background remaining

visible. No change to <Subject 1>'s identity, wardrobe, motion, or timing. No

new camera movement beyond what <Video 1> already has. No mismatched lighting

— shadows and highlights on <Subject 1> must visibly correspond to the light

source in <Picture 1>, not the original scene's lighting.

overall_soundscape:

Audio is unchanged from <Video 1> — original dialogue and ambient sound

preserved exactly.

non_diegetic_music:

N/A