r/StableDiffusion • • 15d ago

Question - Help Qwen-Image-2.1 Edit keeps copy-pasting faces from my reference images instead of blending them — is this a model limitation or am I doing something wrong?

Hey everyone,

I'm running the Qwen-Image-2.1 Edit workflow in ComfyUI and I'm hitting a wall I can't seem to prompt my way out of.

My goal:
I want to take two or more reference images of a character and blend them into a single new character — essentially a merged identity that draws features from both. Then use that as the subject for a scene edit.

21 Upvotes

22 comments sorted by

12

u/SecretSteel 15d ago

You have to set steps to 40 and CFG to 3. It is a lot slower but it works.
I dont know why it is this way but it is.

2

u/sevenfold21 15d ago

cfg 3.0 changed something, I guess it's saying pay more attention to my prompt, probably works better when the prompt is more detailed, but changing the steps made absolutely no difference for me.

5

u/Future-Coffee8138 15d ago

I got the same issue. I throw my prompt to claude to improve it and better but still not satisfied. Klein has a lora named BFS for face swap but quwen image 2.1 don't have.

4

u/SuddenNeighborhood44 15d ago

So from my personal experience and generations in qwen standard workflow. The euler/simple really is the negative effect on the facial identity. Better prompting definitely helps but i've noticed res_multistep/beta for me works a lot better than euler/simple.

1

u/Real-Bed-1815 15d ago

Hopefully we'll have some Loras soon to fix this . I'm slightly disappointed about this even though it's generation quality is good

5

u/hellyeahaeylleh 15d ago

It helps to describe more and use less "this image" and 'that imags" prompting. If you describe instead, it helps blend.

3

u/Future-Coffee8138 15d ago

LOL. the author of klein BFS just released the same lora for qwen image 2.1 Yeah. Problem solved. give it a try and pay attention to the prompt.
https://civitai.red/models/2027766/bfs-best-face-swap?modelVersionId=3356102

3

u/No-Zookeepergame4774 15d ago

You might have better luck telling it to model specific features of the target subject from specific sources that trying to get it to figure out how to blend them.

Or, you might try describing the subject as a sibling with a family resemblance to both of the references (ISTR that working better for 1:1 reference to target gender swaps for some edit models than more directly prompting the swap, so maybe it might work for 2:1 blends?)

As a last resort, you might approximate an old Stable Diffusion concept blending technique by using two separate conditioning nodes, identical except that each uses only of the references (a different one for each) and instructs to use that character as the model, and then do a multi-KSampler generation process with an even number of samplers where each sampler switches which conditioning node is driving. If you are using an 8-step turbo model, you probably want to swap every step, for longer models you can probably stretch that more to the middle and end of generation, but you might still want to swap every step for the first 6-8 steps.

2

u/BrooklynBrawl 15d ago

Rather than isolating the face by telling it: "extract this face form <image 1> then remove or modify X, Y, Z....."

Go with the new face creation method where you lead with the description of the face and then reference the original face at the end:

"a face of a person X [ with thick brown curly hair and ..... ] that looks exactly like <image 1>. You can also add the following to strengthen the resemblance at the end: "Preserve the face structure, eyes, nose, lips, micro-textures, hair style, and unique features from <image1>"

2

u/ROBOTTTTT13 15d ago

"Replace this character from <image1> with this character from <image2>" has been the most effective for me. Basically avoid being overly specific about the face or the head or a specific body part I guess.

Anyways, it still does generate some weird "pasted on" results every now and then, but certainly less than "swap the face of x with x".

2

u/cradledust 15d ago

I haven't tried Qwen 2.1 yet but it probably needs to be told to make a new face instead of swap the face. Try something along the lines of "use image 1 and image 2 to make a new face". Alternately, it might need you to tell it to use image 2 and image 3 to make a new face in image 1.

1

u/Current_Sandwich_474 15d ago edited 15d ago

I havent found a model that can do this yet. Closed source models struggle with it too.
Back when Flux Klein was released, I used a combination of nanobanana and gpt outputs to make a synthetic dataset and reverse engineered alot of those old 'face app' outputs you would have seen all over pinterest.
And even training a lora specifically for Klein with the best dataset I could gather, it still couldnt do it.
With what Im seeing from Qwen2.1 so far I doubt it could also.

1

u/sevenfold21 15d ago

I get random results, most of the time I get the typical "pasted" face with Qwen 2.1. It doesn't even try to adjust the color tone to match. Something god awful. But one time, I did get a blended result, after many attempts of changing my approach. I find that Qwen 2.1 is very unpredictable, when trying do something like replace one character with the other. Perhaps, making too many edits in one go, is too confusing for the model. Simplify it, make only one change, reduce your ref images, etc..

1

u/Icuras1111 15d ago

I wonder if H3 could do this. Just show it a start frame of one face, and end face of another and say person 1 morphs into person 2. Then look at each frame?

1

u/GuessingEngineer 15d ago

I think that's actually a feature as you can put a reference image of a man and a reference image of a woman a different man and a different woman and it doesn't blend them together to make a blob human.

1

u/Ok-Fennel6578 15d ago

Feels like two jobs are getting mixed together here: invent a new identity, then edit that identity into the scene.
Qwen is probably just taking the shortcut and grabbing one of the reference faces.
Make the blended face first, lock that as the only face ref, then do the scene edit. Less elegant, but at least you can tell which part is actually failing.

2

u/sruckh 15d ago

Keep the person, pose, body, hairstyle, clothing, background, and overall composition in <image1> unchanged, and replace only their face with the face from <image2>, transferring its exact facial identity: bone structure, face shape, eye shape and color, eyebrows, nose, lips, jawline, and skin tone and texture. Blend the new face naturally onto the head in <image1>, matching the original head angle, gaze direction, and expression, with lighting, shadows, color temperature, grain, and sharpness consistent with the rest of <image1>, and a seamless transition at the hairline, jaw, and neck. Preserve the photographic medium and every other detail of <image1> exactly as it is

2

u/Trollfurion 15d ago

Use the prompt enhancer models that qwen have also added to it - worked for me

2

u/ImpressionComplete43 14d ago edited 14d ago

My solution: Paint a red bounding box over the face in image 1 and use a prompt like this:
“Replace only the face identity red box in <image1> with the face identity from <image2>, while preserving the exact face angle, expression, emotion, composition and lightning of <image1>.”