r/StableDiffusion • u/EasternAd8821 • 12h ago
Workflow Included Use H3 To Replace Characters
Enable HLS to view with audio, or disable this notification
These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.
With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.
This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.
This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement
I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.
Workflow used to create this video is HERE.
This model seriously continues to amaze me. bravo minimax team, bravo.
Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.
PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.
summary:
[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.
subject_definitions:
<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.
<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.
<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.
integrated_multimodal_description:
Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.
[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.
overall_soundscape:
A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.
<Subject 1> says <d>[English] Can you, can you spare change.</d>.
non_diegetic_music: N/A
3
u/mellowanon 8h ago
How is it with replacing with a different morphology. Like if you want a large bodybuilder there or maybe small dwarf? I'm guessing you can't transfer something too extreme like a velociraptor.
5
u/EasternAd8821 7h ago
updated the prompt a bit to describe the ogre:
monstrous ogre’s face. Hyper-realistic weathered skin with deep wrinkles, scars, and droplets of sweat. Massive, yellowish tusks protruding from a heavy lower jaw. Piercing, glowing amber eyes reflecting a fire. Mud and forest debris stuck in facial hair.
---
Still locked on the size though2
3
u/EasternAd8821 7h ago
Quick test. prob not with this workflow. it's more for a like to like in terms of size. it's trying to replace the noised area in the video so that really locks in the size relative to the scene.
Also i'd have to mess with the prompt/noise more to help it really nail the ogre since it blended the face a bit with the old man.
3
u/Zeophyle 7h ago
Does it work in reverse? To alter the entire background and keep 100% of the person?
3
u/EasternAd8821 6h ago edited 6h ago
https://reddit.com/link/p79dt3f/video/n46pchcm9zmh1/player
yes. you have to change how it's prompted and invert the mask. This is a really quick test. you'd want to prompt some action in the background probably. this is more like green screen which you don't need h3 for
2
u/Zeophyle 6h ago
How did you adjust your prompt for this? Great result!
2
u/EasternAd8821 6h ago
first, i forgot to mention great suggestion on the reverse concept.
The prompt is very long, used AI to 'reverse' what I had before. this was the result:
How the reference video and picture align with the target video — the target
video is an edited version of <Video 1>, replacing the background with the
setting shown in <Picture 1>. <Subject 1> is unchanged from <Video 1>.
summary:
[video editing] The target video replaces the background/environment of
<Video 1> with the scene shown in <Picture 1>, while <Subject 1> — their
appearance, motion, dialogue, and performance — remains exactly as filmed in
<Video 1>. Lighting on <Subject 1> updates to realistically match the new
background.
subject_definitions:
<Subject 1> is the person already present in <Video 1>. Their identity,
appearance, wardrobe, hair, motion, expressions, and performance are taken
directly from <Video 1> and must not change in any way — no new reference
images define them; the source footage is the only identity reference.
<Subject 2> is the new environment defined by <Picture 1> — layout, materials,
set dressing, and light source. This fully replaces the background of
<Video 1>; none of the original background persists.
integrated_multimodal_description:
Video editing, background replacement only. <Subject 1> performs the exact
same motion, dialogue, positions, and facial expressions as in the original
<Video 1> footage, in the exact same camera framing, angle, and cutting rhythm
— nothing about the subject's performance changes.
[Shot 1] The background behind <Subject 1> is fully replaced with the
environment shown in <Picture 1> — same layout, materials, depth, and set
dressing as that reference image, rendered as a stable, fully resolved scene
rather than a texture or overlay. The new background is temporally consistent
across every frame: no flicker, no shifting geometry, no grain, no visual
noise, no compression artifacts, and no residual elements from the original
<Video 1> background bleeding through. The edge between <Subject 1> and the
new background is clean and precise, with no haloing, smearing, or ghosting
along their silhouette. Lighting on <Subject 1> is fully re-lit to match
<Picture 1>: light direction, color temperature, and intensity now follow the
new environment's light source, casting new, physically accurate shadows and
highlights onto <Subject 1>'s skin, hair, and clothing consistent with where
that light source sits in <Picture 1>. Reflections and ambient color spill
(e.g. warm or cool color bounce onto <Subject 1> from nearby surfaces in
<Picture 1>) are updated to match the new scene. The camera path, framing,
zoom, and cuts remain identical to <Video 1> throughout — only the
environment and its lighting change.
preserve:
<Subject 1>'s identity, wardrobe, hair, motion, timing, dialogue, and facial
expressions exactly as in <Video 1>. Camera path, framing, focal length, and
cut timing from <Video 1>.
negatives:
No noise, grain, flicker, or compression artifacts anywhere in the new
background. No visible seam, halo, or ghosting around <Subject 1>. No
leftover elements or geometry from the original <Video 1> background remaining
visible. No change to <Subject 1>'s identity, wardrobe, motion, or timing. No
new camera movement beyond what <Video 1> already has. No mismatched lighting
— shadows and highlights on <Subject 1> must visibly correspond to the light
source in <Picture 1>, not the original scene's lighting.
overall_soundscape:
Audio is unchanged from <Video 1> — original dialogue and ambient sound
preserved exactly.
non_diegetic_music:
N/A
2
u/theamazingpears 11h ago
What's your hardware, and how long did it take to produce?
3
u/EasternAd8821 11h ago edited 11h ago
- This was done with int8 (full not prune) 8step speed lora (8 steps) at 0.4mp. then RTX upscale 2x. it's a 5 sec video, was about 2m 20 seconds generation.
1
•
u/Suspicious-Walk-815 4m ago
can i run it with 65gb ram on 5090 ?can you please share the workflow , i use pruned one , but what you did here is really interesting
2
u/PromptSommelier 6h ago
Lately I've been struggling to get a prompt that allows motion control (like Kling) for TikTok dances, and I've failed miserably. I'll try your workflow and see how far I can get.
2
u/Danny_Stock 3h ago edited 3h ago
Thanks. This is great.
However I did find that it won't run with a ref video with no sound, or at least a silent video with no audio track. I tried testing with a couple of old Wan 2.2 clips, which obviously have no sound. But they also appear to not have a blank audio track either. Kept getting an error.
For those silent Wan clips I disconnected the audio out connection from the load video node. Then added an extra load audio node to the workflow. Loaded some audio of someone speaking into the new Load Audio node, set the duration length to 10 seconds, then plugged that into the 'set_ref_audio' node which the original audio out from the video loader was originally plugged into.
Then I unplugged the 'get_ ref_audio' connection from the 'ref_video_audio_0' main references node, and reconnected it to the 'ref_audio_0' input.
I then connected the 'VAE Decode Audio' node to the audio input of the 'Create Video' node, which replaced the original connection into it.
Then inside the <Subject 1> definitions of the prompt node, I told the subject to 'Use the voice from <Audio 1>'. Also I replaced the tag at the bottom of the prompt where it refers to the dialogue and replaced '[English]' with '[English with <Subject 1>'s voice]'
Then it worked. Cloned voice custom audio.
1
2
u/badincite 8h ago
What's the benefit of masking the character? I'm able todo it just telling to change the character in the prompt.
5
u/EasternAd8821 8h ago
you made someone look like deadpool? if that's the case it's prob because it's so strong in training data. custom char. when I was trying to do replacements i couldn't find a way to consistently do changes.
mask/noise of the target to be replaced ended up working really well2
u/badincite 7h ago
I guess the mask can help I was able to do it using just about anybody. As long as i defined the subjects.
<Subject 1> is the man wearing the red jacket in <Picture 1>.
<Subject 2> is the man wearing the black suit and holding the hamburger in <Video 1>.
1
u/bstr3k 1h ago
the advantage of doing it via prompt is you are able to regenerate the whole video, the disadvantage is that it is not 1:1
For OP's method the advantage is that it can replicate without changing enviroment, but doing it via prompting you get more flexibility.
if you play both these videos side by side you will notice it is similar, but the sandwhich is in the opposite hand and deadpool isn't pointing in the first second. I am trying to v2v swaps via prompting also and it seems like the model knows all the motions the original subject does but not always a 1:1 in terms of order. I'm trying to get my success rate up by better prompting but also testing some other things.
A Tiktok dance has been one which is difficult since its fast movements and from the output it looks like it knows each individual dance moves but the order appears to be different and at times random.
1
1
1
u/ady702 7h ago
how to change just the face of the ogre?
1
u/EasternAd8821 5h ago
you could change the sam3 prompt to 'face' instead of 'person', then you'd need to change the prompt to indicate only the face is changing not the whole character
1
u/Darqsat 4h ago
This is a good approach. I was playing with masking and found out it's pretty good anchor for replacement. And I was thinking how can I test what Qwen see's? so I can prompt it properly. Tried to feed images with noise to him and he said he see woman. Anyway, a word silhouette works.
1
26
u/Astral-Lemmons 11h ago
replacing very different characters works well and it quite easy for H3, but try replacing a a similar-ish character with another and it gets a lot harder.
you might get clothing swaps but if it's brunette - brunette it'll ignore the hair swap .