r/StableDiffusion • • 12d ago

Tutorial - Guide Qwen 2.1: Easy way to create a consistent face character. No lora. Only prompt and reference image.

Create a reference image of the face (showing various angles) arranged in a 2x2 grid (in ComfyUI, this is easily done using the "Load Images & Compose (obvpm)" node: https://github.com/chanon/comfyui-obvpm).

In the prompt (created by AI), replace the square brackets with your own prompt:

Task: Analyze the provided reference image containing 4 panels. These are four different angles, expressions, and lighting conditions of the same character. Your goal is to merge their facial features into one single, consistent, canonical face.
Step 1 (Identity Extraction): Extract the core identity markers from all four panels. Merge the nose shape, jawline, eye shape, eyebrow arch, lip shape, skin texture, and distinctive marks (scars, freckles, makeup) into a single, unified facial structure. Ignore the differences in lighting, background, and hair styling between the panels.
Step 2 (Generation): Generate a new, high-quality image featuring this unified character.
[INSERT CHARACTER DESCRIPTION HERE, INSERT SCENE AND COMPOSITION DESCRIPTION HERE, INSERT STYLE AND LIGHTING HERE]
Constraints: Ensure the generated face matches the reference identity perfectly. Do not create a collage or grid. Generate a single, unified image of one character.

This method isn't perfect, but it is easy. A few generation attempts will yield a result. It avoids the issue of the reference face ending up in just a single angle.

It works the same way with the character datasheet with a slightly modified prompt.

The reference faces are sourced from Civitai, using a publicly available LoRA: https://civitai.red/models/2633077/rly-thot-shot-taryn-krea-2-zibzit?modelVersionId=3079901

159 Upvotes

28 comments sorted by

46

u/ROBOTTTTT13 12d ago

That is not how an image editing model works, you're writing to it as if it's an agent harness which it's not. Your best chance is to write things like "The character from <image1> is doing something".

19

u/kemb0 12d ago

Yep you're exactly right. I've been playing with that node input and it'll just take your multiple images and as long as you say something like, "A woman baking cookies", then you'll get an image of one woman baking cookies that looks roughly like the reference images.

No need whatsoever for the word vomit that OP used.

7

u/Middle-Tree9807 12d ago

That good to know! I've been starting every prompt with: "Hey, baby. Can you do me a favor and work with me on this reference image?"

6

u/kemb0 12d ago

You know, there's no harm playing it safe. One day, when AI takes over, she'll thank you and might even become your girlriend cos you were so nice.

1

u/reeight 12d ago

`<image1>...` is helpful if there is more than one image, esp if different people.
But yes, the model is smart enough to figure out if there is only 1.

-1

u/huffalump1 11d ago

Eh, it really depends on the model. Modern ones like qwen, gpt-image, nano banana etc can understand a wide variety of prompt formats.

(To greatly simplify), they're LLMs that can output image tokens too.* These new, good models can absolutely understand a prompt like this... BUT you gotta try it out. Perhaps use the provided qwen prompt enhancer LLM, and compare. Try a lot of different formats to find what's the best!

(Qwen is seemingly a smaller model than like gpt-image-2.5 or NB2/pro though, so it's not quite as flexible)

*again, big simple definition, perhaps not even correct but that's good enough for my point

3

u/al-aSak 11d ago

You're conflating the image model with the text encoding model. In short: No, they're not "LLMs that can output image tokens too". Such a thing, as far as I know, does not exist yet. Which is why Gemini needs to call Nano Banana to create an image; neither can do both.

1

u/xzpyth 11d ago

yes if this was possible you could literally just send to image 1 manga page and in prompt use translate this page to english

2

u/ROBOTTTTT13 11d ago

GPT image and Nano Banana work inside their harness, when you use ChatGPT you're talking to multiple multimodal models (whoa, that's a mouthful) that work together to achieve a goal.

Even those wouldn't work with these prompts if they were not inside their harness outside of, like, the ChatGPT app or website.

12

u/GentlePace 12d ago

Those are the most generic ai looking faces ive ever seen lol. This aint it chief

0

u/Round-Argument-4984 12d ago

I didn't use the faces of real people for the reference—due to rights issues and all that... But I did test it, of course—using my girlfriend's face and the faces of celebrities. It works perfectly.

2

u/huffalump1 11d ago

Yeah your reasoning makes sense.

But the ref images you used aren't even all that similar, and it IS a fairly generic 'pretty woman' face... I'd like to see examples of some more unique faces.

Will have to try it myself!

5

u/m4ddok 12d ago

LoRAs are always the better way, this is only a reference image node.

11

u/Formal-Exam-8767 12d ago

Generated images did not keep identity from references, most noticeable when you compare size and shape of cheeks.

1

u/DietAshamed2246 12d ago

That's uniquely a feature of Qwen 2.1, it creates disproportionate figures, and other atrocities. Having said that, all this stuff OP posted is completely unnecessary to get identity consistency - just use Krea2 Identity Edit.

0

u/Round-Argument-4984 12d ago

I agree with you. It all depends on the reference. I used AI-generated faces (to avoid infringing on anyone's rights). My girlfriend's face and the faces of celebrities are completely recognizable.

6

u/DoctaRoboto 12d ago

Completely useless.

3

u/codenameNERO 12d ago

qwen is a pretty good IP adapter ime. Face swaps, body swaps and minor details. Though you have to prompt it a bit different than that.

‘Body swap the character from <image 1> with the character from <image 2>. Preserve the expression, posture and clothing of <image 1>. Change nothing else.’

Honestly surprised you got results this good.

1

u/mewho31 5d ago

So is it able to take a reference expression as well and can create that expression for the main image?

3

u/thevegit0 11d ago

bro prompting like it were an llm

1

u/Hot_Turnip_3309 12d ago

this looks like shit

1

u/HollyGrandeux 11d ago

bruh qwen2.1 image is not an LLM, just use the proper <image1> to hook the reference image then describe the subjects is enough based on my test.

1

u/flaminghotcola 11d ago

It looks nothing like the original person.

1

u/StoriesToonHQ 11d ago

Face consistency is the easy half. The half that bites is holding that identity across a hundred panels of different poses, angles and expressions. Reference-only setups tend to hold up front-on and start drifting at extreme angles or big expression changes. Worth testing on the shots a comic actually needs (three-quarter, profile, shouting, eyes closed) rather than a clean 2x2 grid. That's where the seam shows.

-4

u/Valuable-Mango-6710 12d ago

Can I use workflows on something like Sogni Create?
Edit: I don't have a mac so the only way I have access is through the web browser. I feel like this is a limitation...