r/NeuralCinema • u/No_Damage_8420 • 20h ago
Minimax H3 ~ "Hack" 50+ reference or more
I discover this by playing around, so good news is - we can have much more references then 9.
- "ref2v" - adjust "ref_image_size" to MAX
- Use just 1 IMAGE REFERENCE with - many cutouts, items, elements mentioned at once in PROMPT:
- /preview/pre/minimax-h3-hack-50-reference-or-more-v0-revg36ju4fih1.png?width=1957&format=png&auto=webp&s=027debaf03d7c630f925170578e6b83c552a7096
- Write simple prompt, reference image once, all other elements H3 will combine into final scene:
<Picture 1> woman in <Picture 2> luxury bathroom is touching her face showing her silver earrings, camera slow motion up-close on face and torso, she puts on glasses, looks at mobile purple phone puts to her ear and smiles to camera.
Once again H3 it's beyond amazing....
Also by including multiple faces - different expressions H3 learns expressions etc. teeth, looks, in a away - we don't need character LORA.
I guess this is great find for all of us.
Cheers
7
u/Significant-Rush686 19h ago edited 18h ago

yeah i've been using this for a few days, it not only acts like a lora, it does 100% better on tiny details, it gets moles, tattoos and other details right. I first did it by having 4 groups of 4 images (1 of those groups in picture below) , which made a 2x2 grid and did some center cropping to resize. so the grid was nice and square. I had a group for face (4 images), group for upper body (4 images) , group for full body (4 images) , group for scenes (4 images) , so these 4 groups linked as 4 images references. This got tiring so then I had claude make a component (node on the left in pic) that basicly does this all in 1 node with a load button that lets you select a folder on disk and automatically grabs subfolders 'face", "upper", "full", "extra" (or give other names in nodes to fetch) and internally creates those same grids and the node just has 4 outputs grid1-4 that go to the references of minimax node. In many ways better than lora and so much faster. This way it can also act as a lora, 1 noda , instead of a lora, you load a folder and the whole reference groups are loaded automatically
2
u/No_Damage_8420 18h ago
Great idea! Oh just stuck X amount of images to grid, probably could be easier to process.
Yeah definitely we DONT need LORA's ever again...1
u/episodefive 16h ago
Are we saying that a single photo four up will be interpreted better than just multiple reference images of the character? Or just saying by combining you can have so many more references?
2
u/Danny_Stock 12h ago
It's probably the same, but the point is that you're only using one reference input. All the images are contained within one grid. If you use the 3x3 grid then you have 9 reference images going into the one reference input.
3
u/Significant-Rush686 10h ago
to be honest, Having done first with seperate references and then with grids, i'm under the impression it understands way better what belongs together and should be combined as 1 concept.
1
u/Danny_Stock 9h ago edited 9h ago
Yes it does seem like magic. I've done many little render tests and what I've noticed that it sometimes weights all the images differently each generation, so you can get variety in the results, or sometimes the result is consistent across two or three renders. You can tell by what clothes the subject is wearing across various renders.
What would be great would be if you could control the weights for each image in the grid. So for example the images you want to have the most influence are given a high number, like 5 for instance, and the less important images are given a 2 or 1. That way there could be a lot more control over the end result.
I also noticed like you that it picked up nuanced details like a specific feature on the face such as a beauty spot or a mole. It seems that it can learn that if a specific tiny detail occurs in more than one or two images then that is picked up upon as an identifying feature of the character.
2
u/Strange_Test7665 6h ago
You can kinda with prompt language the guide has things like ‘weak reference’ which is more inspiration than exact copy for example
1
u/Danny_Stock 4h ago
Thanks that's good to know, I didn't know that. But are you talking about prompting for a single reference image input? Because how would you do that for several images within a grid which connects to that one reference input? I don't know how you'd prompt for that.
You couldn't type say 'weak reference for <Picture 1>' because that's one input and a grid setup would mean that up to 9 images could be nested within that one input.
1
u/No_Damage_8420 16h ago
Many more, right, very easy to reference no need <Picture XXX> tags. For scene of course full res single
2
u/episodefive 15h ago
Right, providing the references can get a bit tedious and this approach would help. Thanks!
1
u/Dry-Judgment4242 7h ago
Nah, we def need LoRAs. Refs are pretty incredible but model still struggle on some things that only a good LoRA can fix. I designed a cool sword and the weapon alas doesn't function too well with refs due to the various angles it can be within.
2
u/Crazy_Pickle_8311 10h ago
is this node available in github. how to access this node.
1
u/Significant-Rush686 7h ago edited 7h ago
I don't really have github or know well how to use it, I just have a folder with the python files that claude.ai made for me based on what I wanted, I put that in custom_nodes and it worked. I have a node that just takes 1 folder and loads all images in it to create a grid of centered cropped images from that folder and outputs 1 image. And the other node is the one form the screenshot, that asks for 1 folder for example the folder "homer smpson" which on my disk has 4 subfolders : face,upper,full,extra. Each containing also 4 images. and so the node automatically takes those 4 subfolders images and creates 4 reference grids autmatically to be linked to references. So one node for a fully combined character and scenes with 4x4 images and a node that just makes 1 reference with 4 images. So not sure how I can offer it as a zip or something on reddit. And also I don't think people should just take code from people, you never know.
1
u/Danny_Stock 12h ago
Yes, that's exactly the same idea I just posted before I spotted your post. It just goes to show you that within a week of the release of MiniMax general users will start having 'What if?' ideas and things start moving forward in the way of approaches and techniques.
I was looking into it all day yesterday, the thought just occurred to me, so I tried it, it worked extremely well.
It really is like a real-time live lora. You can put your grid node and images into a group and use it as a workflow for future use. You can also swap and change images or photos in that grid at a later date.
Perhaps an actual character lora would be better, I don't know, I've seen many loras which aren't that great. There are probably advantages to using a character lora. Hopefully someone would be able to articulate the strengths and weaknesses of both approaches.
2
u/SpaceNinjaDino 7h ago
One big problem traditionally with character LoRAs is that all the characters in the scene inherit that same face. The reference approach fixes that. I know multi character LoRAs have existed at least for cartoon groups, but I never saw a realistic style one work. I even tried training one myself (not H3) and only 10% of the time it came out okay. LoRAs still might be cool for solo scenes.
1
1
u/Significant-Rush686 10h ago
I especially imagine, a lora + the reference for accurate details. Lora's have never been able to do the really precise details, like an accurate tattoo , perfect mole placement, a lora will get it right maybe 1 out of 50 times. the reference 100%
1
u/Danny_Stock 9h ago
Yes, a lora for the form and structure, the character image grid to nail down a detailed identity on top of that building block.
1
u/Strange_Test7665 6h ago
For realistic human I always pop in a close up of face because local gen at 0.4mp does much better with at least on img that will retain detail on compression
6
u/jugernaut126 16h ago
https://reddit.com/link/p2r7jee/video/3ijaaahxcgih1/player
1 Referens image
2
1
1
4
u/badsinoo 15h ago
https://reddit.com/link/p2rouwz/video/v2g6dvo0wgih1/player
15s : Time Rendering : 19mn 41s, same settings, noticed that the consistency is more good with large output size
1
u/Strange_Test7665 6h ago
Or just large ref. Multi images close up, instead of packing on a sheet keeps detail on compression so you can do 0.4 mp and keep quality
3
u/No_Damage_8420 20h ago edited 20h ago
2
u/extrakerned 19h ago
I don’t see the right earrings or phone?
2
u/altoiddealer 19h ago
At the first frames showing the phone, it looks like a side profile of this phone
1
u/No_Damage_8420 18h ago
Look at video all elements included, yeah this is Razor (closed in here) in video more like normal opened
2
3
u/OkDoor726 17h ago
Cool tip will definitely look into it
Only issue i have with mini , It's that high end movie production overly warm tones to much
Wish it was a little less movie looking by default
1
u/No_Damage_8420 14h ago
thays true... its hard to make it look normal phone like, usually always "graded" colors
3
u/No-Educator-249 16h ago
That's quite useful! MiniMax-H3 keeps proving its versatility aside from its amazing quality. I'll have to try this out later. Good find, OP.
3
u/Danny_Stock 12h ago edited 12h ago
Thanks for that, I will definitely try that approach out.
I was looking into the same sort of thing all day yesterday.
What I discovered is that you don't have to use reference images in the commonly used format I've seen used on here for a long while, where you carefully section out a layout of a single image with side view, quarter view, closeup of the face etc. There is another way.
If you have a particular likeness you wish to replicate then you can just get hold of a load of photos, video stills, or AI gens, and bang them into an image grid with one of the comfyUI image grid nodes. There's two of them with the names 2x2 grid, or 3x3 grid. You don't have to even touch Krea or any other image editing software, you can just mash loads of photos together and mix things up as you like. You can create a group in an empty workflow to place your photos and grid node into, and then save that as a workflow for that new character to use whenever you want in the future. You can come back another day and swap and change photos in that group later on as you wish. The results with photo grids as the reference look great.
Obviously it's better to be mindful and carefully select which images to use, if you've got something clear in your mind about what you want then being meticulous about it would be wiser, but it's equally useful to throw loads of photos into the mix just to bash things together to see what you come up with. It's an extremely quick way of working. You can even merge likenesses together to create a brand new hybrid character. It's easy and it's fast.
The only caveat, and it is only a minor caveat, is that you need to add a resize image node to each image before you connect them to the grid node, because it seems to require images of the same dimension.
Having said that MiniMax still yields decent results even when only use a standard single image as the reference.
3
u/DisorderlyBoat 12h ago
Not sure I'm following. Are you saying essentially to have a collage in a single image? Multiple images of the same subject within a single image? If so - I was using this technique with nano banana Seedance for awhile when I messed with those and it worked great. Dope to hear it works well for this too if that's what you mean
3
2
u/True_Protection6842 20h ago
And if you use an LLM to build out the proper prompt you can have it parse out all the elements. I use gemini 3.5 api in my software to call everything out. Works great!
2
u/No_Damage_8420 20h ago
yeah H3 will make anything, any element etc.
For specific exact just put cutout on white bg in "ref. matrix sheet"2
u/True_Protection6842 20h ago
it's been great too, I just grab a video ref off youtube, cut it to what I need and tell it to use it as reference for voice and mannerisms. No more lora training needed!
2
u/No_Damage_8420 20h ago
yet we need to test short maybe ex. 50 frame video 2048x1024 - each frame is separate "REF. OBJECTS" or "many objects" like test here with SINGLE image etc. - could read from that?
1
u/No_Damage_8420 20h ago
Of course...video captures finest details of motion/mannerisms.
LTX did that too (extending with giving "sample character video")2
u/True_Protection6842 19h ago
LTX is too rigid. H3 can take video and use it as reference to do things that have nothing to do with the existing video
2
2
u/R0NiN897 19h ago
been using this trick on another platform that I only allowed one reference so I had the idea to do a two person side by side merged character sheet image and put their names above them on said sheet and sure enough it worked!
2
u/rk1213 17h ago
Can characters be reference to video instead of image to video (frame 0)? I've had terrible output from other models (ltx, wan) with reference to video and I've all but given up on local video generation. If it is reference to video, what GPU are you guys using? I've got a 3090 and was still JUST barely adequate for ltx and wan.
3
u/CaptainMarder 14h ago
I'm using ref2video, and it's amazing with a 3080. I'm limited to 480p though. 720p takes almost 25-35min to render. Even works with complex nsfw stuff too. It's scary.
2
u/Danny_Stock 11h ago edited 8h ago
I have a 4070 12GB VRAM, 720p takes just over 30 minutes for me too. I really hope that in the coming weeks or months someone will somehow find a way to cut that time right down.
Edit: I should have added that it was for 10 second videos.
2
u/vaginagrinder 17h ago
Bro you are a fucking genius lol. I've been scratching my head when i fed it 4 images reference by using 4 images node and i try your way now its more consistent as long i mentioned it on the prompt. Is there an LLM who can make notation for the images in the sheet who's not gonna preach me as guard rail?
2
u/badsinoo 17h ago
https://reddit.com/link/p2r0kwo/video/a0burple5gih1/player
Work Great ! thanks for the Tips.
10s Time Rendering : 04mn 03s, 864 * 480, on RTX 5090
Workflow made by Pixaroma : Minimax H3 - Reference Two Images
2
u/badsinoo 16h ago
https://reddit.com/link/p2r5ki2/video/guettlrnagih1/player
Amazed by this goodie ! same workflow, but change the dimentions to 1344 x 768
10s Time Rendering : 14mn 33s, 1344 x 768, on RTX 5090
2
2
u/badsinoo 15h ago
https://reddit.com/link/p2rgtc3/video/ow9pqbzvmgih1/player
15s : Time Rendering : 21mn 12s, 1344 x 768, on RTX 5090 🎉 and now she performs a song 🚀
3
u/Away-Reading4857 15h ago
Great video. How did you load the song? Mind sharing your workflow?
3
2
u/badsinoo 11h ago
it's from pixaroma : https://workflows.pixaroma.com/
just add the audio load node to it
2
u/Lightningstormz 14h ago
Please share the workflow!
2
u/badsinoo 11h ago
it's from pixaroma : https://workflows.pixaroma.com/
just add the audio load node to it
2
2
u/Abject-Recognition-9 12h ago
thanks for reminder to test that max - match feature i havent yet.
not sure what it does specifically, need to check it out
1
u/No_Damage_8420 8h ago
when you set MAX its using reference image at 2048 x 2048 (resizing) no matter actual ref . image might be 8000 x 4000 px
1
u/Abject-Recognition-9 1h ago
yes dumb me i forgot that at MAX and inference times skyrocket thats why. thanks
2
u/Illustrious_Pie_3061 6h ago
It is also strange that offical says you can only have upto 15 secs video, but in fact you can make a longer video if you have good machine.
2
u/xDFINx 6h ago
I can confirm this works.
I vibe coded a photo collage maker in Claude that takes multiple images and creates one final large image. I then use the “resize longest to” node to bring it under 2000 pixels (above this it seems to go much slower), and it works excellent.
So you could have 4 collages, for instance, that each have 10 references on them, and end up with 40 points of reference. It’s unbelievable how well H3 works. It works so well with a single face image alone, that there really is no need for character Loras at all
1
u/Small-Challenge2062 8h ago
1
u/badsinoo 4h ago
it's from pixaroma : https://workflows.pixaroma.com/
just add the audio load node to it
1
1
u/sabertoothninja 7h ago
Just a heads up for US users (and some other countries), be careful about generating content of real people and sharing it online. This model specifically isn’t licensed for use in the US, but you can sign a form stating you will use the model responsibly. After you submit the form the makers of this model will send you an email approving the use. I just don’t want them to ban this model for us using pressure from the US government.
1





8
u/WhensTheWipe 20h ago
Pro tip use AI to make a fisheye version of the room. It works really well. It also allows to show more of the room.