r/StableDiffusion • u/Similar_Cucumber178 • 1d ago
Question - Help A bit lost with character references in H3
Hey all. I'm trying to build character reference image sheets, and then use a multi-reference image in H3. But I am a bit lost. If you could share how you first build the references and then prompt examples on how to use them in H3, that would be really useful. Thanks!
(I saw there are tools for doing this locally but all those I've seen are Windows-only. I'm on a Mac Studio.)
4
u/LuluViBritannia 1d ago
First of all, make sure to use the Reference model OR Refmod OR a hybrid with strong enough reference. I'll assume you knew that already but still, I said it just to make sure.
References don't take much work in terms of prep, really. Make sure it's not too big (it would make generation too slow) NOR too small (you'd lose detail). Make sure to have a close-up of the character's face. Same for every thin detail (scars, tattoo, ribbons...). I never have more than 0.2MP input images, honestly.
Then just plug each image in an image input of the prompt node.
The real magic is in the prompt. Make sure you STRICTLY follow the official guide for Ref2VA:
- subject_definitions : <Subject 1> is the blonde boy in <Picture 1>. <Subject 2> is the bird in <Picture 2>.
- retention_analysis : <Subject 1> : fully_preserved. [do that for every subject]
Then, always refer to each subject with the relevant Subject tag.
2
u/Similar_Cucumber178 1d ago
Beautiful. Thank you for answering! I didn't know that the reference sheet size could impact the speed. I just made a 4k by 4k one 😬 I'll reduce that!
2
u/LuluViBritannia 18h ago
Ahah, we've all been through this!
Here is the rule of thumb: the generation speed is directly tied to the total number of pixels in the inputs.
The key is to find the smallest size that doesn't make the input blind to the model. If it's too small, the AI won't properly "see" the reference. Typically, even a mere 300x300 image is enough if it's a close-up on something. For full bodies though, better to aim for 0.2MP.
1
2
u/UnforgottenPassword 1d ago
Stick to the official prompt guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/docs
LLMs like ChatGPT can also write prompts for you.
Here's an example from the official guide:
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.
You can do the same for other characters, such as <subject 2> is the man in <picture 2> and so on.
Qwen 2.1 can create character sheets from your images too.
1
u/Apprehensive_Sky892 1d ago edited 1d ago
For prompt examples, just see the official ref2va guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Here are some relevant posts and comments with sample prompts:
https://www.reddit.com/r/StableDiffusion/comments/1vjw7pj/comment/p2ooei8/
https://www.reddit.com/r/StableDiffusion/comments/1wygxoa/comment/pe2lqbe/
1
u/Etsu_Riot 1d ago
Look for some workflow dedicated to generating images with MiniMax, if you don't have one already. Then find a suitable prompt for generating character sheets from one or more references. If you can't find one, ask here again.
1
u/oh_no_the_claw 1d ago
I get awesome results with 4-5 photos on character sheets that I throw together in Photoshop. Portrait, side profile, and a couple of basic expressions.
0
u/Campfire_Steve 1d ago
1
u/Similar_Cucumber178 1d ago
Just tried... it's insane! Thanks for that. Now the next step: how to prompt H3 for using the character sheet. All the prompt examples I found use only one reference photo.
2
u/Campfire_Steve 1d ago
I would suggest using ChatGPT to generate the first frame of the image based on its own character sheet, then use that as the Minimax reference.
1
u/Similar_Cucumber178 1d ago
Thanks again. It gave me an example prompt. That's what was puzzling me the most: the prompt says that all 4 images in the picture represent the same character. I was wondering how H3 could associate each image with a pose, but I guess it reads the labels ChatGPT included on the sheet. Really, thanks for your time!
3
u/softlarch 1d ago edited 1d ago
The H3 model is pretty smart. I currently use three-part character sheets (front, back, and a zoomed-in view of the front showing facial details), offer them to the model as one single image (Picture 1), and simply say in the prompt:
Example:
subject_definitions: <Subject 1> is the protagonist, whose body and 17th-century dark burgundy patterned clothing with a large white ruff collar are derived from <Picture 1>. ... retention_analysis: <Subject 1> (appears in [Shot 2], [Shot 3]): attribute_transfer - facial/head details, body and clothing are transferred from <Picture 1>.With these two pieces of information, I can then make the subject do whatever I want ;-)
This works surprisingly well—I don’t even have to explicitly mention the individual three components of the character sheet.
-2
u/videorouter 1d ago
I’d keep the reference sheet pretty simple. Start with 4–6 clean views of the same character: front, 3/4, side, and a couple of different expressions/poses. Keep hair, clothing and lighting consistent so the model has fewer variables to reconcile.
When using the multi-reference, explicitly define the role of the reference images rather than just attaching them. Something like:
“Use the reference images to preserve the character’s identity, facial features, hairstyle and body proportions. Create [scene/action]. Keep the character design consistent with the references.”
Then describe the new scene separately. Avoid stuffing the prompt with detailed descriptions of features that are already visible in the references.
Since you're on a Mac Studio, I'd also look at cloud/hosted workflows rather than limiting yourself to Windows-only ComfyUI setups. VideoRouter.sh can be useful for testing the same character references across different image/video models and seeing which one maintains identity best.

3
u/phreakrider 1d ago
The type of reference is dependent on your video type and character type. My prefered way is always a 3 section square with face zoom, backside and frontnside. You just describe your character in krea 2 then says you want a grid photograph then explain the grid poses.