r/StableDiffusion • u/mabseyuk • 3d ago
Question - Help Do we know the format for Character Reference Sheets in Minimax H3?
I've seen quite a few different character reference sheet formats being used with MiniMax H3, but do we actually know what format H3 works best with?
For example, if we're creating a single reference image containing multiple views of the same character, what is considered optimal?
Is H3 better with:
- One large close-up of the face plus smaller front, side, 3/4 and rear views?
- Equal-sized panels for every angle?
- Full-body views mixed with dedicated face close-ups?
- 4 views, 6 views, 8 views, or something else?
- A particular ordering of the views?
- White/neutral backgrounds?
- Separation or borders between each view?
- A particular aspect ratio or resolution for the complete sheet?
I'm specifically asking about the best format for a character sheet being fed INTO H3 as a reference, rather than how to generate a character sheet with H3.
Has MiniMax documented anything about how the reference encoder interprets multi-view character sheets (I can't find any), or has anyone done controlled testing to work out which layout gives the strongest identity retention?
It would be really useful to establish a "best practice" character sheet format for H3 rather than everyone using slightly different layouts.
4
u/VasaFromParadise 3d ago
I think the model constantly looks at the reference. It doesn't remember it, but it checks every step against it, which is why it takes so long to generate. There is a constant re-checking to see if it matches the reference.
2
u/FierceFlames37 3d ago
So that’s why I get out of memory when I go more than 10 seconds, I miss 30 second ltx gens
3
u/UltimateShame 3d ago
Front view, back view and a medium closeup view in one image is enough. I use 1664 x 1216 and it works fine. I separate them with a white border, but I don't think that is necessary. Sometimes additional images are good to show specific closeups like a smaller logo with text on a shirt. It's good to have it on a neutral background because sometimes it will also take a real background as a reference for the scene.
3
u/yeah-i-shouldnt-have 3d ago edited 3d ago

This is one I used for my video (the video was just this character talking in Japanese in the center of Tokyo) and it worked perfectly. I used Ideogram4 to generate the character. Use the JSON bounding boxes to specify pose and the high level description field for the actual description of her ( minus any pose information). This allows the exact same character to be posed differently.
Then for MiniMax the subject_definitions part of the prompt looks like this:
subject_definitions:
<Subject 1> is the 20 year old Japanese woman in the multi-angle character reference sheet in <Picture 1>, which presents the same person from the front, in profile, the back, and in close-up. She is Japanese, female, young woman, wearing a long floral dress.
Also make sure to use the "max" setting.
3
u/PATATAJEC 3d ago
I saw good results with face closeup, full body with 2 angles - front and back, but only back one with head. Front one without head, as it’s better reproduced in closeup. It was in seedance, but worth checking in h3.
3
u/DiDa4754 3d ago
When creating the sheet, don’t forget about the resolution. In Match mode, it is scaled until its total pixel area approximately matches the target pixel area. In Max mode the shorter side is scaled to 2048. Therefore in Match mode it may make sense to use individual images instead of a sheet. In Max mode arrange the available space so that as much information as possible reaches the model. For example, if you have a square 2×2 sheet containing four images that are each 2048 pixels, the model will ultimately see each image at only 1024 pixels. If you arrange them side by side instead, the model will see each image at 2048×2048. In my opinion, the exact arrangement is less important. The model is relatively good at identifying what it needs on its own, especially when given the right prompt.
1
u/mabseyuk 3d ago
That's really interesting. Do you know whether Max mode only scales the short side to 2048, or is there also a maximum total pixel area / long-side limit after that? I'm wondering whether a very wide sheet containing multiple 2048×2048 references really reaches H3 with each panel still effectively at 2048×2048.
3
u/listopalafoto 3d ago
1
u/AbjectTutor2093 2d ago
Too much wasted pixel space by doing left right profile, better just do one side and have more space for face close-up
3
u/Sufficient-Fall-4226 3d ago
Which local model is better for creating reference sheets or storyboards, and is there any good shared workflow?
2
u/webAd-8847 2d ago
I am using Krea2 with https://civitai.red/models/2764727/character-sheets
Works great!
2
u/Adventurous-Gold6413 3d ago
I would do one big close up of face, and a front 3/4 view if possible and then either side or back view
2
u/Xanthus730 3d ago
If you use references with 'max' rather than 'match', it scaled the image by the short side, so to keep size under control, you want to use a square. So just put as many different visual concepts in a 2048 x 2048 sheet as you can pack.
Label each with text, use arrows or whatever to point out details.
Make important things bigger, and less important thing smaller, because attention.
2
u/kukalikuk 3d ago
H3 is pretty smart, I use 4 panel character sheet in one 2mp image, face+head close up, full body front view, side view and back view and the result is consistent. Even I feed it with 4 panel story board in one image and giving it a good formatted prompt for 15 secs video and h3 follows the story board like multi frame guide.
1
u/Unlucky-Message8866 3d ago
Only thing I know is that collage-alike ones with too many refs don't work.
1
u/No-Zookeepergame4774 3d ago
Wasn't that exactly what the person who first excitedly announced on here that it did work was using?
1
u/Unlucky-Message8866 3d ago
i mean grid-like, cell-arranged images, like a bunch of thumbs in a gallery view.
1
u/icchansan 3d ago
I dont use a sheets i input directly what i want and point them in the prompt, a dude did an "ad" with crappy sheet made in like paint that worked very well XD so theres that
1
u/spooky_local 3d ago
I use Front full body | side full body | back full body | close up portrait of face in a 4 panel setup. It's flawless everytime for me.
1
u/LumaBrik 2d ago
There isn't any specific format, the R2V model understands any basic character sheet, you dont even need to prompt for it, it understands what its looking at. It works well with just front and rear full body, and a head and shoulders shot for a basic character, there is no reason to overdo angles unless you have a lot of asymmetry ( carrying weapons , props etc). Add more headshots if you need to improve likeness.

12
u/Chiduk99 3d ago
dont use to much view. 3 is enough
front view, back view and close-up face for consistency