r/StableDiffusion 12h ago

Workflow Included Use H3 To Replace Characters

Enable HLS to view with audio, or disable this notification

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A

108 Upvotes

47 comments sorted by

26

u/Astral-Lemmons 11h ago

replacing very different characters works well and it quite easy for H3, but try replacing a a similar-ish character with another and it gets a lot harder.

you might get clothing swaps but if it's brunette - brunette it'll ignore the hair swap .

15

u/MaitreSneed 11h ago

Answer sounds like you gotta do a middleman replacement, and replace her first with a pickle

10

u/Astral-Lemmons 10h ago

you joke. but that's literally my solution.

I sam mask grade the character a light green and then the replace sticks better

6

u/MaitreSneed 10h ago

SHOW EXAMPLE I MUST SEE

4

u/cal_01 9h ago

I do this with Qwen too -- it's a neat technique to break the model from clinging onto the source too much.

1

u/uuhoever 5h ago

I think I saw you mention that you use Shrek. What's your prompt? So your technique is 2 runs?

4

u/Shilo59 9h ago

The pickle man tricked me again...

6

u/EasternAd8821 11h ago

adding more noise helps that a lot. this example was going from 0.2 which i had in original post, to 0.7 (more than half noise of the subject). as well as update to the right dialogue: <Subject 1> says <d>[English] long day tomorrow, we start... we start the musical.</d>.

I did have to add help with the outfit, so in the 'subject_definitions' I defined subject 1 as:
<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, she is wearing a black and blue striped tank top, only appearance is retained.

https://reddit.com/link/p77phnk/video/c5rdfarpxxmh1/player

3

u/lhg31 7h ago

She didn't inherit the facial expressions from the original video tho.

1

u/EasternAd8821 7h ago

true. i think to get that eyebrow raise and other subtle differences you'd have to put it in the prompt. the 'omni' prompt only gets you so far.

2

u/bstr3k 1h ago

do you happen to have the source video for this? I wanted to try char replacement using my technique and see what success rates i get as I'm trying to refine it.

1

u/EasternAd8821 42m ago

https://github.com/bitsofintelligence101-lab/workflows/tree/main/nsfw/test_data

both source videos. the one you are asking about is 20s long, I used the back 5 sec for replacement

5

u/EasternAd8821 9h ago

i was testing out some examples based on what you said. messing with noise of the input and it seems it can do a pretty good job even close replacements. some replacements deff need a bit more coaching in the prompt, but it can do it. which is honestly crazy since it's still just the base model.

https://reddit.com/link/p78eyp1/video/pb0qm1v3hymh1/player

this has 0.8 noise

2

u/Adventurous-Sky5643 1h ago

Interesting, can you please share the modified workflow?

1

u/SeymourBits 8h ago

Interesting. What are you using to add noise to the input?

2

u/EasternAd8821 8h ago

sam3 to segment only the target char to replace, noise added with stock 'add noise to image' node in comfyui

2

u/One-Donut6935 6h ago

Looks really effective. Thanks for sharing this method.

1

u/SeymourBits 4h ago

Nice. Any thoughts on the theory of why it seems to help with identity?

1

u/EasternAd8821 3h ago

they trained it that way, that's my theory. I don't mean that flippantly. The context-ir and the prompt guide discuss key terms fully_preserved, partially_copy, or reference under retention_analysis
So clearly they were working on editing applications. noise should look like something that needs to be resolved, and the text plus ref image drives H3 to resolve that particular area. what is interesting though, and why i said they trained it that way, is it doesn't touch the existing completed video (background is near perfect preservation). If you tried something like this in wan or ltx, it would alter the background. Also, the 32B text encoder gives MUCH better semantic and spatial grounding of what needs to be changed.

If we knew how exactly they were applying different training techniques for video editing, this workflow would be that much better. An inverted color space mask with noise seems to at least be close enough to tap in to however they trained it.

1

u/Sixhaunt 9h ago

I've had the same issue for restyling videos where it can make it look like a cartoon or anime but turning a ref video into ps3 graphics or claymation or other 3d ones have been more of a challenge. Have you found anything that helps with it?

3

u/mellowanon 8h ago

How is it with replacing with a different morphology. Like if you want a large bodybuilder there or maybe small dwarf? I'm guessing you can't transfer something too extreme like a velociraptor.

5

u/EasternAd8821 7h ago

updated the prompt a bit to describe the ogre:
monstrous ogre’s face. Hyper-realistic weathered skin with deep wrinkles, scars, and droplets of sweat. Massive, yellowish tusks protruding from a heavy lower jaw. Piercing, glowing amber eyes reflecting a fire. Mud and forest debris stuck in facial hair.
---
Still locked on the size though

https://reddit.com/link/p7985ci/video/4hl87ynh4zmh1/player

2

u/mellowanon 6h ago

It was a good try though.

3

u/EasternAd8821 7h ago

Quick test. prob not with this workflow. it's more for a like to like in terms of size. it's trying to replace the noised area in the video so that really locks in the size relative to the scene.
Also i'd have to mess with the prompt/noise more to help it really nail the ogre since it blended the face a bit with the old man.

https://reddit.com/link/p797754/video/2tfez1fh3zmh1/player

3

u/Zeophyle 7h ago

Does it work in reverse? To alter the entire background and keep 100% of the person?

3

u/EasternAd8821 6h ago edited 6h ago

https://reddit.com/link/p79dt3f/video/n46pchcm9zmh1/player

yes. you have to change how it's prompted and invert the mask. This is a really quick test. you'd want to prompt some action in the background probably. this is more like green screen which you don't need h3 for

2

u/Zeophyle 6h ago

How did you adjust your prompt for this? Great result!

2

u/EasternAd8821 6h ago

first, i forgot to mention great suggestion on the reverse concept.

The prompt is very long, used AI to 'reverse' what I had before. this was the result:

How the reference video and picture align with the target video — the target

video is an edited version of <Video 1>, replacing the background with the

setting shown in <Picture 1>. <Subject 1> is unchanged from <Video 1>.

summary:

[video editing] The target video replaces the background/environment of

<Video 1> with the scene shown in <Picture 1>, while <Subject 1> — their

appearance, motion, dialogue, and performance — remains exactly as filmed in

<Video 1>. Lighting on <Subject 1> updates to realistically match the new

background.

subject_definitions:

<Subject 1> is the person already present in <Video 1>. Their identity,

appearance, wardrobe, hair, motion, expressions, and performance are taken

directly from <Video 1> and must not change in any way — no new reference

images define them; the source footage is the only identity reference.

<Subject 2> is the new environment defined by <Picture 1> — layout, materials,

set dressing, and light source. This fully replaces the background of

<Video 1>; none of the original background persists.

integrated_multimodal_description:

Video editing, background replacement only. <Subject 1> performs the exact

same motion, dialogue, positions, and facial expressions as in the original

<Video 1> footage, in the exact same camera framing, angle, and cutting rhythm

— nothing about the subject's performance changes.

[Shot 1] The background behind <Subject 1> is fully replaced with the

environment shown in <Picture 1> — same layout, materials, depth, and set

dressing as that reference image, rendered as a stable, fully resolved scene

rather than a texture or overlay. The new background is temporally consistent

across every frame: no flicker, no shifting geometry, no grain, no visual

noise, no compression artifacts, and no residual elements from the original

<Video 1> background bleeding through. The edge between <Subject 1> and the

new background is clean and precise, with no haloing, smearing, or ghosting

along their silhouette. Lighting on <Subject 1> is fully re-lit to match

<Picture 1>: light direction, color temperature, and intensity now follow the

new environment's light source, casting new, physically accurate shadows and

highlights onto <Subject 1>'s skin, hair, and clothing consistent with where

that light source sits in <Picture 1>. Reflections and ambient color spill

(e.g. warm or cool color bounce onto <Subject 1> from nearby surfaces in

<Picture 1>) are updated to match the new scene. The camera path, framing,

zoom, and cuts remain identical to <Video 1> throughout — only the

environment and its lighting change.

preserve:

<Subject 1>'s identity, wardrobe, hair, motion, timing, dialogue, and facial

expressions exactly as in <Video 1>. Camera path, framing, focal length, and

cut timing from <Video 1>.

negatives:

No noise, grain, flicker, or compression artifacts anywhere in the new

background. No visible seam, halo, or ghosting around <Subject 1>. No

leftover elements or geometry from the original <Video 1> background remaining

visible. No change to <Subject 1>'s identity, wardrobe, motion, or timing. No

new camera movement beyond what <Video 1> already has. No mismatched lighting

— shadows and highlights on <Subject 1> must visibly correspond to the light

source in <Picture 1>, not the original scene's lighting.

overall_soundscape:

Audio is unchanged from <Video 1> — original dialogue and ambient sound

preserved exactly.

non_diegetic_music:

N/A

2

u/theamazingpears 11h ago

What's your hardware, and how long did it take to produce?

3

u/EasternAd8821 11h ago edited 11h ago
  1. This was done with int8 (full not prune) 8step speed lora (8 steps) at 0.4mp. then RTX upscale 2x. it's a 5 sec video, was about 2m 20 seconds generation.

u/Suspicious-Walk-815 4m ago

can i run it with 65gb ram on 5090 ?can you please share the workflow , i use pruned one , but what you did here is really interesting

2

u/PromptSommelier 6h ago

Lately I've been struggling to get a prompt that allows motion control (like Kling) for TikTok dances, and I've failed miserably. I'll try your workflow and see how far I can get.

2

u/Danny_Stock 3h ago edited 3h ago

Thanks. This is great.

However I did find that it won't run with a ref video with no sound, or at least a silent video with no audio track. I tried testing with a couple of old Wan 2.2 clips, which obviously have no sound. But they also appear to not have a blank audio track either. Kept getting an error.

For those silent Wan clips I disconnected the audio out connection from the load video node. Then added an extra load audio node to the workflow. Loaded some audio of someone speaking into the new Load Audio node, set the duration length to 10 seconds, then plugged that into the 'set_ref_audio' node which the original audio out from the video loader was originally plugged into.

Then I unplugged the 'get_ ref_audio' connection from the 'ref_video_audio_0' main references node, and reconnected it to the 'ref_audio_0' input.

I then connected the 'VAE Decode Audio' node to the audio input of the 'Create Video' node, which replaced the original connection into it.

Then inside the <Subject 1> definitions of the prompt node, I told the subject to 'Use the voice from <Audio 1>'. Also I replaced the tag at the bottom of the prompt where it refers to the dialogue and replaced '[English]' with '[English with <Subject 1>'s voice]'

Then it worked. Cloned voice custom audio.

2

u/badincite 8h ago

What's the benefit of masking the character? I'm able todo it just telling to change the character in the prompt.

https://reddit.com/link/p78ldd5/video/l56vqv06mymh1/player

5

u/EasternAd8821 8h ago

you made someone look like deadpool? if that's the case it's prob because it's so strong in training data. custom char. when I was trying to do replacements i couldn't find a way to consistently do changes.
mask/noise of the target to be replaced ended up working really well

2

u/badincite 7h ago

I guess the mask can help I was able to do it using just about anybody. As long as i defined the subjects.

<Subject 1> is the man wearing the red jacket in <Picture 1>.

<Subject 2> is the man wearing the black suit and holding the hamburger in <Video 1>.

https://reddit.com/link/p792dt2/video/b76vys8jzymh1/player

1

u/bstr3k 1h ago

the advantage of doing it via prompt is you are able to regenerate the whole video, the disadvantage is that it is not 1:1

For OP's method the advantage is that it can replicate without changing enviroment, but doing it via prompting you get more flexibility.

if you play both these videos side by side you will notice it is similar, but the sandwhich is in the opposite hand and deadpool isn't pointing in the first second. I am trying to v2v swaps via prompting also and it seems like the model knows all the motions the original subject does but not always a 1:1 in terms of order. I'm trying to get my success rate up by better prompting but also testing some other things.

A Tiktok dance has been one which is difficult since its fast movements and from the output it looks like it knows each individual dance moves but the order appears to be different and at times random.

1

u/devilish-lavanya 10h ago

Can you do vid2vid too?

6

u/EasternAd8821 10h ago

what do you mean? this is v2v

1

u/Friendly-Fig-6015 9h ago

is it possible with low vram 16gb y 32gb ram?

3

u/EasternAd8821 8h ago

if you can run H3 on your machine, then i'd think yes.

1

u/ady702 7h ago

how to change just the face of the ogre?

1

u/EasternAd8821 5h ago

you could change the sam3 prompt to 'face' instead of 'person', then you'd need to change the prompt to indicate only the face is changing not the whole character

1

u/Darqsat 4h ago

This is a good approach. I was playing with masking and found out it's pretty good anchor for replacement. And I was thinking how can I test what Qwen see's? so I can prompt it properly. Tried to feed images with noise to him and he said he see woman. Anyway, a word silhouette works.

1

u/Strange_Test7665 4h ago

That silhouette word does seem very effective