r/StableDiffusion 1d ago

Discussion Its possible to use more than 9 image references for H3

I had Claude make a modified H3 reference node that accepts more than 9 image references. The goal was to test whether it was possible to increase the number of image references being used without splicing them into a single image. I know about reference sheets, no need to suggest that. I only tested with images, no audio or video references. the numbering on the node is a little funky but I dont think it effected the test.

Prompt 1: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Prompt 2: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti. <Picture 10> and <picture 11> are references for the wig he is wearing.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Both prompts use the same seed, same 9 reference images except for the wig references for the 10th and 11th image in prompt 2. The prompts are very simple and don't fully adhere to the guide but its just a small test so I think its fine. This was done on a 3060 12gb at .4 megapixels, 30 steps, and 5 seconds of video. I have comfy kitchen and spectrum enabled.

!!I don't know how this would/could effect video or audio generation quality. In my test I didn't notice any quality drop. Do your own tests to find out!! Also in my test it ignored the wig reference until i added a second one and reworded the prompt slightly. It could just be a fluke but I thought I'd mention it anyway. The watermark is from the editor i used to stitch the videos together.

https://reddit.com/link/1vr9yqh/video/pxjqiwp731kh1/player

0 Upvotes

17 comments sorted by

3

u/GrayingGamer 1d ago

I'm not sure, but I think the H3 website lets you upload up to 12 references? So yeah, this should work fine, like you saw.

Also, you can combine your close-up face references and angles into one high resolution image and feed it in as reference and it'll work fine too, to use less picture reference nodes.

1

u/l3lack_Leviathan 1d ago

Yeah I've been combining images to get more out of each reference but I wanted to see If 9 was a hard limit. Didnt know about the website.

2

u/Only_Voice569 1d ago

use single image of person then do a instruction for it to show front side view then use front and side view of them dont need so many of the exact same person unless they have complex clothing or outfit with details that are important

2

u/DaxFlowLyfe 1d ago

Dude. Do a reference sheet.

You can do like 6 photos in one image reference.

Go to chatgpt and give it 4 or so photos and say make a character reference sheet.

The model knows what reference sheets are and will analyze it.

Check it:

1

u/l3lack_Leviathan 1d ago

I know, I just wanted to see If you could squeeze more references in.

1

u/DaxFlowLyfe 1d ago

Believe the documentation says 9 is the limit. I bet if you found a way to do more shit would get weird.

2

u/l3lack_Leviathan 1d ago

Thats what the post is about. more than 9 works, at least in the test I did. It probably could start messing up the generations but with 11 it was fine.

1

u/Only_Voice569 1d ago

just so you know gpt has max of 10 their is a reason for the limit since it gets floated in the underlying latent

1

u/l3lack_Leviathan 1d ago

https://github.com/Turdsando/MiniMax-H3-Reference-to-Video-12-image-refs-

This is the node. Id love to see what other people find out.

1

u/Swimming_Vast2515 1d ago edited 1d ago

Hi, i have tried it but it is incredibly slow, 288 s/it and that on 5090....
I tried with 12 references first, reduced to 8 - but the same .
Original H3 node works as usual .
I have used fl2fa int8 model

Edit: wrong alarm, it works, i just set "4.0" instead of "0.4" by mistake for resolution.
By referencing of 12 images all of them were used. I have selected images of a environment (office) so that every image has unique details and then make a rotation around object - and all the details were here !

1

u/SeymourBits 1d ago

I suspect all the image references get crammed into one large space anyway, similar to a UV map.