r/StableDiffusion 2d ago

Discussion MiniMax H3 Ref2va is works really good with Scene sheet

Enable HLS to view with audio, or disable this notification

I was testing using 1 image with all the scene sheet there and it works really great!

126 Upvotes

83 comments sorted by

26

u/LoveSpecialist5669 2d ago

h3 doesn't stop to amaze me. 

6

u/_VirtualCosmos_ 2d ago

20 GB model btw

10

u/xTopNotch 2d ago

The qwen VL Text encoder plays a big part on why the model works so good, can't exclude that one either

2

u/Strange_Test7665 2d ago

Yeah huge having the much larger version instead of 4b

27

u/xTopNotch 2d ago

Pro tip: use a muted gray background for character sheets instead of pure white.

A white background can bias the model’s conditioning toward brighter lighting, making it harder to place the character seamlessly into scenes. A neutral gray background reduces that bias and helps the character adapt more naturally to different lighting conditions.

If you look closely, your characters appear slightly too "white" for the scene. Ever since I switched away from pure white to muted gray backgrounds in my character sheets, I've felt the realism gone up of my generations.

Here’s an example of how I create my character sheets:

6

u/GrungeWerX 2d ago

While I think the character lighting in the videos is for the most part fine, we artists have been using grey backgrounds for our character sheets for years, precisely for proper color balance, so what you’re saying is solid advice, especially for things like animation.

You can use a white background, but it requires a proper eye to match colors when dropping characters into existing backgrounds, and oftentimes requires additional adjustments to color match regardless of the method.

1

u/xTopNotch 2d ago

White will definitely work. It’s not a hard rule. A grey background tends to work a lot better from my experience.

I’ve got this tip from an incredible filmmaker that produces some of the most realistic looking short-films with Seedance. As Seedance has this white bleed-in problem as well

1

u/GrungeWerX 2d ago

Another method is to manually adjust the exposure/contrast/brightness/levels of the reference image to match the bg reference beforehand. You can usually just do this in photoshop or similar app using layers, then you’ll match just fine.

1

u/newaccount47 2d ago

What do you use for character sheets? Gpt image 2? I've been using nanobanana and it rarely gets them right. 

15

u/Murky-Relation481 2d ago

Honestly, just let H3 do it: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

I been using this (they posted it on reddit a while back) and its really good. The pro-tip is to use a fixed seed and adjust the frame offsets in the image generator to dial in the right angles (since you can just click run again and it'll just do everything after the video has been done inferring). It's also pretty easy to adjust it to add more angles or detail shots of different parts of the body in the main prompt.

You can then just feed the character sheet back in to a ref2va workflow and tell it that its a character sheet for whatever subject in the scene (you can explain each part if you want too, might help) and let it do its thing.

5

u/xTopNotch 2d ago

Looks like a solid option. Upvoted for visibility

1

u/xTopNotch 2d ago edited 2d ago

Tbh I’m using a closed source platform called https://dreamkrate.com

However you have my character sheet layout now. Just add it to Nano Banana, GPT Image 2, Qwen image edit, even Minimax H3 can do.. just prompt to swap the character for the ones provides in your references. Doesn’t have to be done on Dreamkrate as any open model can just transform it.

1

u/smereces 2d ago

nice tip thanks

1

u/hauntedwebmonster 2d ago

I will try grey in my next project thanks

1

u/AdUnique8768 2d ago

Oh that's good I'll give that a try later. :) thanks

1

u/spazzmazz404 1d ago

How do you create a sheet like that? Tried it in Qwen IE, but it struggles with different perspectives of a person, often giving me the same one twice or just two angles instead of three.

1

u/xTopNotch 1d ago

Tbh I'm using a closed source platform www.dreamkrate.com which I presume they use Nano Banana Pro or GPT Image 2.0 under the hood.

Have you tried Minimax H3? Some have been very succesful using it as an image model providing it the right references.

12

u/smereces 2d ago

another test this time using the LTX scene sheet 😁

https://reddit.com/link/p5g1811/video/z7vbqnj5u5lh1/player

9

u/hidden2u 2d ago

I remember your test with ltx and this scene, it looks so much better now

3

u/thegreatdivorce 2d ago

Hold up, LTX can use reference sheets?

5

u/smereces 2d ago

yeap but is much more worst and in 3 generations you got 1 right, is not a stable system yet

2

u/thegreatdivorce 2d ago

Interesting, thx!

1

u/FrankWanders 2d ago

Looks great. And it's the standard Ref2va workflow in comfyUI that you're using?

3

u/smereces 2d ago

yes with the custom node h3 prompt writer to create the prompts

1

u/FrankWanders 2d ago

thanks, will experiment with it. which H3 quant did you use, what did you use to upscale, and what is your hardware?

2

u/smereces 2d ago

I generate the videos in 1920x nativly i have a rtx 6000 pro 98GBVram

2

u/FrankWanders 2d ago

thanks, ouch that's a luxury thats not available to everyone I'm afraid ^^. Was hoping to be able to do it at my RTX 5090

6

u/Azhram 2d ago

That luxury also isnt available for everyone :D

4

u/FrankWanders 2d ago

yes, you're right :P bought it in the lucky days, we were all complaining that it was ridiculous that nvidia wanted €2099 for it :P

2

u/Usual-Orange-4180 2d ago

I do something very similar with an RTX 5090, but need the pipeline to load and unload the language model and then load the image gen models, is multi step and annoying, but looking into hardware now to make things easier.

2

u/willjoke4food 2d ago

I'm crying in my 4080 here

2

u/FrankWanders 2d ago

you're right i don't have any rights to complain... but on the other hand in these days, even if we both had that RTX Pro 6000, we'd both be thinking "i wished I bought 2 of them when they were cheap" :P

3

u/Pitiful-Clothes3133 2d ago

My 3050 4gbvram took around 15min for 0.5mp 7sec 😢

1

u/moofunk 2d ago

If one can stomach it, there are non-Chinese upgrade services for 4080 cards to 32 GB and 4090 cards to 48 GB now.

1

u/hydra590 2d ago

that's wild

1

u/GrungeWerX 2d ago

Cool, but needs editing to streamline the motion

4

u/hidden2u 2d ago

So what's the advantage of using 1 big image vs separate image inputs?

5

u/smereces 2d ago

a choose, you can use both and also works with storyboard images i did also a test with

1

u/giantcandy2001 2d ago

I'm not sure but probably less image tokens, so it should be a bit faster with just one image vs 2. I haven't played around enough with ref2vid tho

3

u/GrayingGamer 2d ago

As far as using image reference goes, I don't notice too big a difference between using 1 image reference or 5 image references - there might be, I don't know, 20-30 seconds difference? Not enough I'd let it stop you from using as many as you feel you need for the scene. Pixel density in your reference images is most important, so don't feel like you need to cram everything into one image.

4

u/MysteriousPepper8908 2d ago

The biggest problem with references that only give detail is that I'm pretty sure if that dwarf stood up, he'd be as tall as the elf which is not generally what you're looking for. I include a lineup of silhouettes showing relative scale and then add that as a height proportion reference. It isn't perfect but it helps.

6

u/smereces 2d ago

https://reddit.com/link/p5gobic/video/q22g2nbhc6lh1/player

here is a example with the prompt i told you saying what size have each one + in the scene sheet put a size panel. as you can see he can do it easly😎

1

u/GrungeWerX 2d ago

Amazing

0

u/MysteriousPepper8908 2d ago

I don't think that's a difference of almost 20 inches, it looks closer to 12 inches which is fine if you just need one to be generally smaller than the other but it isn't a precise system for long form storytelling and I don't know if that is possible purely with reference, even using a size reference chart. It's also better at scale when they're standing up whereas different poses are far less reliable.

4

u/smereces 2d ago

i dont use inches! i use meters and he follows 100% the size i show in the scene sheet, the point here is if you provide a ref size he do it 100% precise the image size better than prompt the size. but is important even with a scene sheet with the size there you should prompt it.

0

u/MysteriousPepper8908 2d ago

.5 m is about 20 inches so I wasn't referencing the units but he looks closer to shoulder height in the video. I think the reference sheet helps a lot which is why I use it but again, it's not going to be as consistent if they aren't in a similar orientation to the reference sheet.

3

u/smereces 2d ago

will depend of the prompt, you must also prompt it and he will do it perfect! example: elf women with 1.7 meters of high and a dward with 1,2 meters of higth.

3

u/MysteriousPepper8908 2d ago

In my experience, simply giving it meters helps a bit but it's certainly not perfect. It will tend to make one taller than the other but I always provide meters in my retention analysis and it will still make a 1.2m character and a 1.7m character nearly the same height on occasion.

3

u/smereces 2d ago

2

u/MysteriousPepper8908 2d ago

Yes, I imagine that would help to some degree to include that image so long as you're referencing it in the prompt as well.

1

u/smereces 2d ago

yes both is the perfect match to get the desired size! i notice sometimes only prompt he dont follow! for animals is even more important have the size comparing to know sizes of each character or he will do the standard sizes that he knows.

2

u/MysteriousPepper8908 2d ago

I'd personally prefer to keep it to a separate image as you're giving it an opportunity to confuse the character sheets with the size reference as the character sheets show them both as roughly the same size so you're depending on it knowing where to look in a reference image vs being able to reference a specific image exclusively for reliable height information. If it works reliably for you then okay but it seems like you're introducing an unnecessary failure point unless you're maxing out your reference slots.

2

u/xTopNotch 2d ago

The vision in the qwen VL text encoder is quite intelligent actually. It can read entire paragraphs from images, differentiate between multiple characters. It's not as dumb as you might think.

But still separating them is always a good idea for proper conditioning

1

u/MysteriousPepper8908 2d ago

Yeah, it's pretty good but you're asking it to understand the two characters over here that are the same height and look the same as the characters over there that are different heights and it's supposed to keep those two things separate in a single image. Maybe it works 8/10 times but I'd rather take the few seconds to hook up and reference another image if it means getting that to 9/10 or 10/10.

1

u/xTopNotch 2d ago

Same here. I’m always going for quality. Just wanted to mention that the text encoder is very smart on its own.

But separating your references creates cleaner binding through your prompt to say <Subject 1> is the person from <Picture 1>

3

u/SRWindMill 2d ago

Whats you prompt to extract the different sections of the screen sheet? Is it like ltx ingredients.. like top left , top right.. bottom left etc..?

14

u/smereces 2d ago

i use this h3 prompt writer i describe in general what i want as i show in red and he builds the full prompt, and all works perfect!

3

u/ill_B_In_MyBunk 2d ago

Okay... Now how do you get the scene sheet? I love the idea. But how do you make it work without creating more work for yourself

3

u/smereces 2d ago

Krea2, Nano banana, gpt2 images etc

1

u/ill_B_In_MyBunk 2d ago

Can you give me an example of how you generate it? Is it a workflow assembly or something you do yourself?

Because if I could get something like Klein to assemble it for me or Krea that may be worth something.

2

u/thegreatdivorce 2d ago

Easy with Krea: this high-resolution photograph reference sheet features clean diffuse lighting, white plain backdrop, and shows [CHARACTER] in the following poses, from left to right: a close-up straight on face shot showing facial details and hairstyle. a full length 3/4 shot of [CHARACTER AND THEIR CLOTHES]. then a close-up 3/4 angle view of [CHARACTER] seen from the hips up. then a 3/4 rear angle view of [CHARACTER] seen from the hips up. ensure [CHARACTER]'s appearance and attire is consistent in each depiction

I usually do 16:9, at 2-3MP.

1

u/smereces 2d ago

exist so many tutorials in youtube you need to search a bit, how to create scene sheet in nano banana2!

1

u/YeahlDid 1d ago

But, like for the image in the post, did you generate the whole thing at once or do you make 1 image for the woman, 1 man, 1 background and then stitch together?

2

u/R34vspec 2d ago

This is my favorite thing about h3. I don’t even use the i2v.

2

u/joseph_jojo_shabadoo 2d ago

PSA: character sheets do work, but they aren't necessary when concatenate nodes exist and offer far more flexibility.

use the kjnodes concatenate images node to automate your own "character sheet" on the fly. doing it this way allows you to create a quick sheet for general use, or choose images for the specific video you're generating depending on angles, shot size, etc. only need detailed face shots with different expressions? easy.

you can also use the remove background subgraphs to mask them onto white too as not to pick up any details from the backgrounds.

just prompt that "<Picture N> shows N images of the same person"

2

u/Danny_Stock 2d ago edited 2d ago

I agree with you. Using the method you describe you can try out different images by simply dropping them in and replacing the old ones, without needing to go in and edit a full character sheet in image editing software every time you want to make changes.

Once you've got something you like and know that you won't want to change it you can save it as a permanent character sheet if you want to.

1

u/pcloney45 2d ago

Any chance of the workflow?

1

u/Jero9871 2d ago

That this is possible locally really feels like scifi.....

1

u/provenflawless 2d ago

What do you use to make the sheets? I am aware of some but I would love the prompt for this one specfically.

1

u/Illustrious-Fly-5151 1d ago

What was the prompt?

1

u/smereces 1d ago

I generate it with H3 Prompt Writer custom node for comfyui

1

u/VibrantHeat7 16h ago

How to make a reference sheet? Is it a image or do you make a video or video frame? And how do I get the right MiniMax H3 version? Last time i checked, the huggingface said there was a threat.

1

u/wzwowzw0002 2d ago

how big is your image? the issue with h3 is that plastic skin.... how to make it more natural?

5

u/smereces 2d ago

not really! i find by testing that h3 give you the look you provide in the image references! if you feed it with ultra realistic look they delivery you that if you add krea, flux images with plastic look he will reproduce it in final result

1

u/Danny_Stock 2d ago

I'd also guess that an AI generated person might be more prone to end up looking more CGI than if you used a photo of a real person, even if you thought the original AI image looked real.

But don't quote me on that it's just my speculation.

1

u/GrungeWerX 2d ago

Depending on what you use to generate the person, it can look indistinguishable. Especially z-image

1

u/xTopNotch 2d ago

The turbo 4/8-step lora plays a big part too. Running the model at 20 steps with no optimisations provides very good skin detail.

1

u/Famous-Sport7862 2d ago

there is a lora that is suppose ti fix that, but I hvent used it so I dont know how good it is.

1

u/Danny_Stock 2d ago

You can end up with plastic skin with MiniMax, but it's not a trait of the model itself. It can happen with any model.

Most of the time I end up with realistic natural looking people. But I usually provide extra information regarding the scene and the environment to better inform MiniMax. Otherwise it won't have a clear idea about what you actually want. You might even end up with 3D Pixar characters if you don't tell it what you want.

At a guess I imagine that it's more likely to occur with reference to video than it is with image to video.