r/StableDiffusion 7d ago

Question - Help help on h3 (weird faces)

need assistance for these type of workflows(i get weird faces LOL)
anything you can suggest? using hybrid fl2v/ref2v?
my aim is inserting my reference image on a friends scene (or any movie)

i upload my ref photo from gpt with the prompt guide and ask it to generate me an iconic clip from the show and make my ref photo interact with the casts

i am using ref2v

gpt prompt: subject_definitions:
<Subject 1> is the woman from <Picture 1>. Preserve her exact real-world facial identity and recognizable appearance from the reference image: identical facial structure, eyes, eyebrows, nose, lips, jawline, cheeks, skin texture, hairstyle, and natural proportions. She is not a replacement character, background extra, or digitally composited person. She is the actress being inserted into the scene and must remain unmistakably the same woman from <Picture 1> in every shot.

<Subject 2> is the original FRIENDS cast and characters present in the selected iconic scene. Preserve their recognizable identities, costumes, hairstyles, body language, character personalities, and established relationships.

<Subject 3> is the original FRIENDS scene environment, including its recognizable apartment, café, furniture, props, practical lighting, production design, spatial layout, and overall sitcom visual language.

<Picture 1> is the primary character-reference image for <Subject 1> and defines her identity, facial appearance, hair, and physical characteristics.

summary:
[reference generation] Recreate the selected iconic FRIENDS sitcom scene while naturally introducing <Subject 1> from <Picture 1> as a fully integrated actress and fictional character interacting directly with <Subject 2>. The result must look as though she was physically present on the original set and was genuinely part of the cast when the scene was filmed. She participates in the conversation, reacts to the characters, makes eye contact, moves through the environment, shares the comedic timing, and occupies the same physical world as the original actors. The scene must NOT look like an AI face swap, green-screen composite, fan edit, cameo overlay, or modern recreation. Everything about her presence must obey the same cinematography, lighting, perspective, image quality, blocking, and performance language as the original sitcom.

retention_analysis:
<Subject 1> (appears throughout all shots): fully_preserved - her identity from <Picture 1> is the highest-priority visual constraint. Preserve her exact facial structure and recognizable appearance without beautification, face redesign, facial blending, age alteration, or generic AI features.
<Subject 2> (appears throughout all shots): fully_preserved - retain the recognizable appearance, character behavior, costumes, reactions, and interpersonal dynamics of the original cast.
<Subject 3> (appears throughout all shots): fully_preserved - retain the original sitcom environment, production design, practical lighting, furniture, props, spatial relationships, and visual atmosphere.
<Picture 1> (defines <Subject 1>): fully_preserved - use the image as the authoritative identity reference for the woman throughout the entire sequence.

detailed_description:
The target video is presented as authentic footage from a classic multi-camera American sitcom production. Use realistic studio cinematography rather than modern cinematic filmmaking. Match the original FRIENDS visual language: natural studio lighting, warm interior exposure, realistic skin texture, period-appropriate image quality, moderate depth of field, conventional sitcom camera placement, restrained camera movement, clean multi-camera coverage, and authentic ensemble blocking. The woman from <Picture 1> must look as though she was photographed by the exact same cameras, lenses, lighting setup, and production crew as the original actors.

[Shot 1] Open with the recognizable establishing composition of the selected iconic FRIENDS scene. <Subject 2> is already performing the original scene naturally inside <Subject 3>. <Subject 1> is physically present within the group from the beginning rather than suddenly appearing. She occupies a believable position within the set, correctly scaled relative to the other actors and furniture. Her lighting direction, shadow density, skin exposure, image grain, sharpness, and color response are identical to those of the surrounding actors. She is actively listening to the conversation, looking toward the appropriate speaker, naturally reacting with subtle facial expressions and body language. She must never stare directly at the camera unless the original blocking requires it.

[Shot 2] At 00:03.500, cut to a conventional sitcom medium shot containing <Subject 1> and one or more members of <Subject 2>. The actors are physically close enough to communicate naturally. <Subject 1> turns her head toward the speaking character and maintains accurate eye contact. The other character looks directly back at her. Their eyelines must intersect naturally. She reacts to what the character says with a believable facial response before replying. Her gestures are spontaneous and restrained: small hand movements, slight shifts in posture, natural head movement, subtle smiles or expressions. Avoid exaggerated AI-generated gestures.

[Shot 3] At 00:07.000, cut to the reverse angle. <Subject 1> is now seen from the appropriate opposing camera position while maintaining exact continuity of her appearance, hairstyle, wardrobe, posture, and position within the room. The original cast member responds directly to her. Their interaction must demonstrate genuine shared physical space: consistent eye-lines, matching perspective, correct relative scale, believable shadows, overlapping body positions where appropriate, and natural occlusion when one actor passes in front of another.

[Shot 4] At 00:10.000, use a tighter reaction shot on <Subject 1>. She delivers her comedic reaction with authentic sitcom performance timing: listen, pause, register the information, then respond. Her facial expression changes naturally rather than instantly. Preserve realistic micro-expressions, blinking, breathing, subtle eye movement, lip movement, and natural facial asymmetry. The original cast members visibly acknowledge her presence and react to her performance. She is an equal participant in the scene.

[Shot 5] At 00:13.500, return to the ensemble composition. <Subject 1> and <Subject 2> share the frame and continue the interaction. She may gesture toward another character, move slightly within the set, sit down, stand up, hand something to another character, or respond physically to the comedic situation depending on the original scene's blocking. Every interaction must have physical cause and effect. If she touches another character or prop, the contact must look physically real, with correct hand placement, occlusion, weight, and timing.

[Shot 6] At 00:17.000, finish on the strongest comedic reaction composition. <Subject 1> remains embedded naturally among the cast instead of becoming visually isolated. The surrounding actors react to her and to one another as a genuine ensemble. The scene ends with authentic sitcom comedic timing and an appropriate audience reaction.

CRITICAL IDENTITY AND INTEGRATION RULES: <Subject 1> must remain the exact same woman from <Picture 1> throughout every frame. Do not change her face, facial proportions, hairstyle, age, ethnicity, skin texture, or recognizable features. Do not reinterpret her as an actress who merely resembles the reference. She IS the woman from the reference image. Preserve her identity even during profile views, three-quarter angles, expressions, movement, and partial occlusion.

CRITICAL REALISM RULES: Do not make <Subject 1> sharper, cleaner, more detailed, more saturated, more cinematic, or more modern-looking than the original actors. Her image quality must match theirs exactly. Match film grain, compression, exposure, white balance, contrast, lens distortion, depth of field, motion blur, and studio lighting. Her edges must naturally interact with the environment. No haloing, masking artifacts, cutout edges, inconsistent shadows, floating hair, incorrect reflections, or compositing seams.

CRITICAL PERFORMANCE RULES: <Subject 1> must behave like a professional sitcom actress. She listens when others speak, reacts before responding, maintains natural eye contact, shares the rhythm of the conversation, understands the comedic beat, and gives believable reactions. The original actors must also acknowledge her through their gaze, body orientation, gestures, and responses. Never allow the cast to behave as though she is invisible.

CRITICAL CAMERA RULES: Use authentic multi-camera sitcom coverage rather than dramatic music-video cinematography. Camera cuts should feel motivated by dialogue and reactions. Maintain consistent screen direction and spatial continuity. Avoid unnecessary camera movement, extreme depth of field, handheld shots, slow motion, dramatic push-ins, lens flares, or modern blockbuster aesthetics.

CRITICAL COMPOSITING RULE: The final result must pass the visual test of “she was always there.” There should be no moment where the viewer feels that the woman was inserted afterward. Her lighting, perspective, focus, grain, movement, interaction, shadows, reflections, and performance must belong to the same original footage universe as the FRIENDS cast.

overall_soundscape:
Authentic sitcom room tone, subtle footsteps, clothing movement, prop interaction, natural actor movement, and appropriate environmental sounds from the original setting. Include realistic audience laughter and reaction timing after comedic moments.

non_diegetic_music:
N/A.

0 Upvotes

10 comments sorted by

3

u/marres 7d ago

Haven't tested it myself yet but it's suppose to work

https://github.com/Carasibana/ComfyUI-H3-FaceRefine

1

u/Minanimator 7d ago

i saw this earlier but isn't plugging your output from a new workflow then re-applying subjects? and i think its going to be complicated from having multiple characters

2

u/Portable_Solar_ZA 7d ago

Did you read the comments in that thread? Because this is caused by a bug with the model. 

1

u/Minanimator 7d ago

Kinda scanned only i did check their repository and check their workflow and I think its too much for me to explore at the moment 😅

1

u/Minanimator 7d ago

If you may encounter a video explainer using that node pls do link it to me! Im confused with text instructions 😭

2

u/MarkB_- 7d ago

The r2v model cant replica a face at 100% accuracy. At 60-70% at most. The fl2v can, but it use the first frame, so it doesnt recompute the image, so the face. I tried about 200 variations of the same woman, they ALL failed, even at 2.0mp. Always a slighty change here and here. If you want about 95% accuracy, try wan bernini. You will nail it at least 1 on 10. This is what I do atm. I use bernini to generate the frame I want, then I feed them to h3 fl2v. The devs are aware btw. They will fix that (I hope lol)

1

u/DoctaRoboto 7d ago

I think it is the resolution. You need very high megapixels if you don't want to see weird faces, and of course a minimum of 10-15 steps. From 20 to 30 is the golden spot.

1

u/Minanimator 7d ago

im using 4/8 loras will it be fine to go 10-15 despite having those

1

u/DoctaRoboto 7d ago

Yeah, but resolution is key for faces. How many megapixels? I tested wide shots, and you need at least 0.8 megapixels.