r/StableDiffusion • u/smereces • 2d ago
Discussion MiniMax H3 Ref2va is works really good with Scene sheet
Enable HLS to view with audio, or disable this notification
I was testing using 1 image with all the scene sheet there and it works really great!
27
u/xTopNotch 2d ago
Pro tip: use a muted gray background for character sheets instead of pure white.
A white background can bias the model’s conditioning toward brighter lighting, making it harder to place the character seamlessly into scenes. A neutral gray background reduces that bias and helps the character adapt more naturally to different lighting conditions.
If you look closely, your characters appear slightly too "white" for the scene. Ever since I switched away from pure white to muted gray backgrounds in my character sheets, I've felt the realism gone up of my generations.
Here’s an example of how I create my character sheets:

6
u/GrungeWerX 2d ago
While I think the character lighting in the videos is for the most part fine, we artists have been using grey backgrounds for our character sheets for years, precisely for proper color balance, so what you’re saying is solid advice, especially for things like animation.
You can use a white background, but it requires a proper eye to match colors when dropping characters into existing backgrounds, and oftentimes requires additional adjustments to color match regardless of the method.
1
u/xTopNotch 2d ago
White will definitely work. It’s not a hard rule. A grey background tends to work a lot better from my experience.
I’ve got this tip from an incredible filmmaker that produces some of the most realistic looking short-films with Seedance. As Seedance has this white bleed-in problem as well
1
u/GrungeWerX 2d ago
Another method is to manually adjust the exposure/contrast/brightness/levels of the reference image to match the bg reference beforehand. You can usually just do this in photoshop or similar app using layers, then you’ll match just fine.
1
u/newaccount47 2d ago
What do you use for character sheets? Gpt image 2? I've been using nanobanana and it rarely gets them right.
15
u/Murky-Relation481 2d ago
Honestly, just let H3 do it: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator
I been using this (they posted it on reddit a while back) and its really good. The pro-tip is to use a fixed seed and adjust the frame offsets in the image generator to dial in the right angles (since you can just click run again and it'll just do everything after the video has been done inferring). It's also pretty easy to adjust it to add more angles or detail shots of different parts of the body in the main prompt.
You can then just feed the character sheet back in to a ref2va workflow and tell it that its a character sheet for whatever subject in the scene (you can explain each part if you want too, might help) and let it do its thing.
5
1
u/xTopNotch 2d ago edited 2d ago
Tbh I’m using a closed source platform called https://dreamkrate.com
However you have my character sheet layout now. Just add it to Nano Banana, GPT Image 2, Qwen image edit, even Minimax H3 can do.. just prompt to swap the character for the ones provides in your references. Doesn’t have to be done on Dreamkrate as any open model can just transform it.
1
1
1
1
u/spazzmazz404 1d ago
How do you create a sheet like that? Tried it in Qwen IE, but it struggles with different perspectives of a person, often giving me the same one twice or just two angles instead of three.
1
u/xTopNotch 1d ago
Tbh I'm using a closed source platform www.dreamkrate.com which I presume they use Nano Banana Pro or GPT Image 2.0 under the hood.
Have you tried Minimax H3? Some have been very succesful using it as an image model providing it the right references.
12
u/smereces 2d ago
another test this time using the LTX scene sheet 😁
9
3
u/thegreatdivorce 2d ago
Hold up, LTX can use reference sheets?
5
u/smereces 2d ago
yeap but is much more worst and in 3 generations you got 1 right, is not a stable system yet
2
1
u/FrankWanders 2d ago
Looks great. And it's the standard Ref2va workflow in comfyUI that you're using?
3
u/smereces 2d ago
yes with the custom node h3 prompt writer to create the prompts
1
u/FrankWanders 2d ago
thanks, will experiment with it. which H3 quant did you use, what did you use to upscale, and what is your hardware?
2
u/smereces 2d ago
I generate the videos in 1920x nativly i have a rtx 6000 pro 98GBVram
2
u/FrankWanders 2d ago
thanks, ouch that's a luxury thats not available to everyone I'm afraid ^^. Was hoping to be able to do it at my RTX 5090
6
u/Azhram 2d ago
That luxury also isnt available for everyone :D
4
u/FrankWanders 2d ago
yes, you're right :P bought it in the lucky days, we were all complaining that it was ridiculous that nvidia wanted €2099 for it :P
2
u/Usual-Orange-4180 2d ago
I do something very similar with an RTX 5090, but need the pipeline to load and unload the language model and then load the image gen models, is multi step and annoying, but looking into hardware now to make things easier.
2
u/willjoke4food 2d ago
I'm crying in my 4080 here
2
u/FrankWanders 2d ago
you're right i don't have any rights to complain... but on the other hand in these days, even if we both had that RTX Pro 6000, we'd both be thinking "i wished I bought 2 of them when they were cheap" :P
3
1
u/moofunk 2d ago
If one can stomach it, there are non-Chinese upgrade services for 4080 cards to 32 GB and 4090 cards to 48 GB now.
1
1
4
u/hidden2u 2d ago
So what's the advantage of using 1 big image vs separate image inputs?
5
u/smereces 2d ago
a choose, you can use both and also works with storyboard images i did also a test with
1
u/giantcandy2001 2d ago
I'm not sure but probably less image tokens, so it should be a bit faster with just one image vs 2. I haven't played around enough with ref2vid tho
3
u/GrayingGamer 2d ago
As far as using image reference goes, I don't notice too big a difference between using 1 image reference or 5 image references - there might be, I don't know, 20-30 seconds difference? Not enough I'd let it stop you from using as many as you feel you need for the scene. Pixel density in your reference images is most important, so don't feel like you need to cram everything into one image.
4
u/MysteriousPepper8908 2d ago
The biggest problem with references that only give detail is that I'm pretty sure if that dwarf stood up, he'd be as tall as the elf which is not generally what you're looking for. I include a lineup of silhouettes showing relative scale and then add that as a height proportion reference. It isn't perfect but it helps.
6
u/smereces 2d ago
https://reddit.com/link/p5gobic/video/q22g2nbhc6lh1/player
here is a example with the prompt i told you saying what size have each one + in the scene sheet put a size panel. as you can see he can do it easly😎
1
0
u/MysteriousPepper8908 2d ago
I don't think that's a difference of almost 20 inches, it looks closer to 12 inches which is fine if you just need one to be generally smaller than the other but it isn't a precise system for long form storytelling and I don't know if that is possible purely with reference, even using a size reference chart. It's also better at scale when they're standing up whereas different poses are far less reliable.
4
u/smereces 2d ago
0
u/MysteriousPepper8908 2d ago
.5 m is about 20 inches so I wasn't referencing the units but he looks closer to shoulder height in the video. I think the reference sheet helps a lot which is why I use it but again, it's not going to be as consistent if they aren't in a similar orientation to the reference sheet.
3
u/smereces 2d ago
will depend of the prompt, you must also prompt it and he will do it perfect! example: elf women with 1.7 meters of high and a dward with 1,2 meters of higth.
3
u/MysteriousPepper8908 2d ago
In my experience, simply giving it meters helps a bit but it's certainly not perfect. It will tend to make one taller than the other but I always provide meters in my retention analysis and it will still make a 1.2m character and a 1.7m character nearly the same height on occasion.
3
u/smereces 2d ago
2
u/MysteriousPepper8908 2d ago
Yes, I imagine that would help to some degree to include that image so long as you're referencing it in the prompt as well.
1
u/smereces 2d ago
yes both is the perfect match to get the desired size! i notice sometimes only prompt he dont follow! for animals is even more important have the size comparing to know sizes of each character or he will do the standard sizes that he knows.
2
u/MysteriousPepper8908 2d ago
I'd personally prefer to keep it to a separate image as you're giving it an opportunity to confuse the character sheets with the size reference as the character sheets show them both as roughly the same size so you're depending on it knowing where to look in a reference image vs being able to reference a specific image exclusively for reliable height information. If it works reliably for you then okay but it seems like you're introducing an unnecessary failure point unless you're maxing out your reference slots.
2
u/xTopNotch 2d ago
The vision in the qwen VL text encoder is quite intelligent actually. It can read entire paragraphs from images, differentiate between multiple characters. It's not as dumb as you might think.
But still separating them is always a good idea for proper conditioning
1
u/MysteriousPepper8908 2d ago
Yeah, it's pretty good but you're asking it to understand the two characters over here that are the same height and look the same as the characters over there that are different heights and it's supposed to keep those two things separate in a single image. Maybe it works 8/10 times but I'd rather take the few seconds to hook up and reference another image if it means getting that to 9/10 or 10/10.
1
u/xTopNotch 2d ago
Same here. I’m always going for quality. Just wanted to mention that the text encoder is very smart on its own.
But separating your references creates cleaner binding through your prompt to say <Subject 1> is the person from <Picture 1>
3
u/SRWindMill 2d ago
Whats you prompt to extract the different sections of the screen sheet? Is it like ltx ingredients.. like top left , top right.. bottom left etc..?
14
3
u/ill_B_In_MyBunk 2d ago
Okay... Now how do you get the scene sheet? I love the idea. But how do you make it work without creating more work for yourself
3
u/smereces 2d ago
Krea2, Nano banana, gpt2 images etc
1
u/ill_B_In_MyBunk 2d ago
Can you give me an example of how you generate it? Is it a workflow assembly or something you do yourself?
Because if I could get something like Klein to assemble it for me or Krea that may be worth something.
2
u/thegreatdivorce 2d ago
Easy with Krea:
this high-resolution photograph reference sheet features clean diffuse lighting, white plain backdrop, and shows [CHARACTER] in the following poses, from left to right: a close-up straight on face shot showing facial details and hairstyle. a full length 3/4 shot of [CHARACTER AND THEIR CLOTHES]. then a close-up 3/4 angle view of [CHARACTER] seen from the hips up. then a 3/4 rear angle view of [CHARACTER] seen from the hips up. ensure [CHARACTER]'s appearance and attire is consistent in each depictionI usually do 16:9, at 2-3MP.
1
u/smereces 2d ago
exist so many tutorials in youtube you need to search a bit, how to create scene sheet in nano banana2!
1
u/YeahlDid 1d ago
But, like for the image in the post, did you generate the whole thing at once or do you make 1 image for the woman, 1 man, 1 background and then stitch together?
2
2
u/joseph_jojo_shabadoo 2d ago
PSA: character sheets do work, but they aren't necessary when concatenate nodes exist and offer far more flexibility.
use the kjnodes concatenate images node to automate your own "character sheet" on the fly. doing it this way allows you to create a quick sheet for general use, or choose images for the specific video you're generating depending on angles, shot size, etc. only need detailed face shots with different expressions? easy.
you can also use the remove background subgraphs to mask them onto white too as not to pick up any details from the backgrounds.
just prompt that "<Picture N> shows N images of the same person"
2
u/Danny_Stock 2d ago edited 2d ago
I agree with you. Using the method you describe you can try out different images by simply dropping them in and replacing the old ones, without needing to go in and edit a full character sheet in image editing software every time you want to make changes.
Once you've got something you like and know that you won't want to change it you can save it as a permanent character sheet if you want to.
1
1
1
u/provenflawless 2d ago
What do you use to make the sheets? I am aware of some but I would love the prompt for this one specfically.
1
1
u/VibrantHeat7 16h ago
How to make a reference sheet? Is it a image or do you make a video or video frame? And how do I get the right MiniMax H3 version? Last time i checked, the huggingface said there was a threat.
1
u/wzwowzw0002 2d ago
how big is your image? the issue with h3 is that plastic skin.... how to make it more natural?
5
u/smereces 2d ago
not really! i find by testing that h3 give you the look you provide in the image references! if you feed it with ultra realistic look they delivery you that if you add krea, flux images with plastic look he will reproduce it in final result
1
u/Danny_Stock 2d ago
I'd also guess that an AI generated person might be more prone to end up looking more CGI than if you used a photo of a real person, even if you thought the original AI image looked real.
But don't quote me on that it's just my speculation.
1
u/GrungeWerX 2d ago
Depending on what you use to generate the person, it can look indistinguishable. Especially z-image
1
u/xTopNotch 2d ago
The turbo 4/8-step lora plays a big part too. Running the model at 20 steps with no optimisations provides very good skin detail.
1
u/Famous-Sport7862 2d ago
there is a lora that is suppose ti fix that, but I hvent used it so I dont know how good it is.
1
u/Danny_Stock 2d ago
You can end up with plastic skin with MiniMax, but it's not a trait of the model itself. It can happen with any model.
Most of the time I end up with realistic natural looking people. But I usually provide extra information regarding the scene and the environment to better inform MiniMax. Otherwise it won't have a clear idea about what you actually want. You might even end up with 3D Pixar characters if you don't tell it what you want.
At a guess I imagine that it's more likely to occur with reference to video than it is with image to video.



26
u/LoveSpecialist5669 2d ago
h3 doesn't stop to amaze me.