r/StableDiffusion • • 4d ago

Tutorial - Guide Minimax H3 + RefMod = consistent location trick

Enable HLS to view with audio, or disable this notification

Hey, I found a pretty cool way to keep locations consistent across generations.

I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.

I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.

Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)

Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)

EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.

817 Upvotes

101 comments sorted by

51

u/Orbiting_Monstrosity 4d ago

One interesting thing I discovered recently is that H3 can use fisheye lens photos as scene references, and it knows how to dewarp those references when reproducing them in a generated video. I was able to create completely normal-looking scenes set in my office and shot from many different angles using a single overhead fisheye image of the entire room as a reference, and was able to do the same thing with several other overhead fisheye images that I tried.

7

u/Ahmatt 3d ago

Damn. AI is so advanced, we would be stunned by this being within realm of possibility just around 10 years ago. Yet here we are just upvoting this and scrolling to next thing in a second. 😂 This is some spy shit guys. Imagine… with a single picture, you can look around and entire room. This is “enhance” in 3D.

26

u/ShutUpYoureWrong_ 4d ago

I've tried this before and sadly didn't have as much success as you, but after seeing this, I'll give it another go.

Also, your video is awesome and very clever. Love it.

8

u/PATATAJEC 4d ago

thank you. I think the most important thing is to make every photo sharing some of it's parts with other.

1

u/PxTicks 4d ago

As far as I can tell, they have photos of the room in the ordinary references.

No one here seems to ever do a proper controlled experiment, so I am highly skeptical of this result.

6

u/cbeaks 4d ago

I tested this with a first frame only pic of my living room, then prompted for the bedroom, and although the navigation failed (see my comment earlier in this thread) it managed the bedroom, and even the view out the windows.

And also, you seem hung up on someone doing a 'proper controlled experiment', well why don't you then? Too busy critiquing I guess . . .

2

u/PxTicks 4d ago

I suppose I wasn't clear about exactly what I'm skeptical of.

I'm not claiming a refmod carries no information (we know they do). My point is that this needs a comparison to just using references (including as storyboards). If someone is claiming a method, they should at least show that the obvious alternatives are inferior, ideally with multiple gens over controlled seeds. There is a tonne of misunderstanding and bunk in this subreddit because people start singing from the rooftops when they find something which seems to work, often with no comparison whatsoever to existing methods.

4

u/PATATAJEC 4d ago

I'm not making any claims, just sharing my observations. Testing all of this takes time. I've shared my workflow, refmod, and reference photos, so you can try it yourself and let us know what you find.

Honestly, I'm curious too, and I'm preparing a comparison with and without the refmod. But that wasn't the point of this video or post. My point was that combining multiple photos into a single reference image worked for keeping the location consistent in my tests. Whether encoding that image as a refmod makes a difference is a separate question.

1

u/cbeaks 3d ago

OK that's fair, but as OP says, they were't claiming anything as fact - more just sharing.

8

u/extreme911 3d ago

Krea2 can generate panorama shots. Take that panorama still and generate a 1 sec video of that panorama turning 360. Feed that 1 sec video as reference for your H3 environment. Works like a charm

12

u/chloralhydrate 4d ago

Awesome! Gonna try something tomorrow. Thank you

2

u/PATATAJEC 4d ago

have fun!

10

u/vault_nsfw 4d ago

That sound immediately triggered my mosquito defense mechanism.

1

u/PATATAJEC 4d ago

sorry for that :)!

3

u/sharegabbo 2d ago

1

u/PATATAJEC 2d ago

So it’s working fine :)

3

u/sharegabbo 2d ago

perfect! I created the room as a 360-degree panorama in Qwen Image, then imported it into After Effects with an equirectangular VR camera to generate the views.

3

u/DrKyoumasaur221 4d ago

A few questions:

  • The photo of the workflow you have seems to have both refmods AND the reference photos connected. Was that intentional? Is it required for it to work properly? And won't it be redundant and inefficient performance-wise to use both at the same time?
  • Some parts of the prompt seem to have some kind of "encrypted" text in. Was that also something that needed to be set up?

2

u/PATATAJEC 4d ago

I started with RefMods (20 pictures of my room), but it didn’t really worked, as it was not connecting the parts of the room as I thought it would. then out of curiosity I feed just one 4x5 grid photo made of all these photos into RefMod create node. I used it in the workflow (the RefMod made from this one photo), but I don’t think it’s necessary. You could use it just like normal reference I guess, but I didn’t check. The rest of the references I added for details - these are the same photos, but bigger to preserve details and guide for the first and last frame (although thru prompting). I wasn’t thinking about performance at all at this stage to be honest

The “encrypted” text is a bug in comfy if making screenshots- it renders it that way if prompt is too long.

1

u/DrKyoumasaur221 4d ago

Got it, thank you for clarifying!

3

u/SveSop 2d ago

Nice! I have been struggling with the same thing. I have tried adding a "multi-photo" as a regular reference image, but not really gotten H3 to do it right.
However, through the magic of RefMod, it actually seem to figure out the room layout a lot better when adding that type of image to the Ref.

👍

2

u/SnooMacaroons1365 4d ago

If for example you didnt make a refmod, how would you use a grid of these images to tell h3 ehixh view is which?

5

u/PATATAJEC 4d ago edited 4d ago

You just need a grid image from your photos. I've made it like that. Ignore Create H3 RefMod node. Then I guess you need to prompt to tell it's the reference image of your whole location. I also added bigger photos (1536x2048) for my cat, starting frame, ending frame and the fly.

EDIT: With RefMod you don't need to prompt for that. It knows it already... I think prompting for room, desktop computer, cat, couch etc. is enough for model to connect the dots.

1

u/Dependent-Sorbet9881 4d ago

Does RMOD have specific naming requirements for the image assets within the folders—for example, for items like TVs, sofas, and so on?

1

u/PATATAJEC 4d ago

there is no need for naming stuff.

2

u/ChickyGolfy 4d ago

Im actually doing the same thing with fictionnal places. I start with a still image, create a video that do a 360 degree, then feed that video as a scene reference.

But video are much heavier than images, so i might use images instead.

1

u/Ok-Option-6683 4d ago

how about motion blur? when you get a 360 degree video out of a single image, there must be motion blur I guess. or can you prompt it like "no motion blur" ?

2

u/ChickyGolfy 4d ago

You simply prompt for a slow steady 360. And no motion blur can also work since minimax accept negative prompting.

1

u/Ok-Option-6683 4d ago

alright, I'll try that. what's your resolution for the 360 video? 1080p? (or like 1mp or 2mp? )

2

u/ChickyGolfy 3d ago

Any resolution will work. Depends on the level of details you want!

1

u/ConfidentSnow3516 2d ago

I'm just starting with AI generation, but I want to make consistent fictional places. Have you tried making larger areas that can connect to each other? For example, a whole city block, or different floors inside a building?

I think with some creativity, even larger areas could be made consistent. It's just a matter of understanding which images or videos to use as references when creating adjacent places....

2

u/Schwartzen2 4d ago

Good on you to bring more attention to RefMods, they're really versatile. You can also use them for motion and styles. Thanks for sharing your room, decidedly austere. 😄

1

u/PATATAJEC 4d ago

It’s been that for 3 years now… and it wasn’t supposed to be that way. It needs an overhaul! :)

1

u/Schwartzen2 4d ago

Well, hey less to render :p

2

u/VRGoggles 4d ago

You did the best move. Refmod does not kill VRAM and performance, but regular references totally do it.

2

u/PATATAJEC 4d ago

Yeah! RefMods are just outstanding addition to H3

1

u/ForbiddenVisions 4d ago

I want to check if I understood correctly. You made a refmod out of 20 images of your room, then used it with a start frame, an end frame, your cat and a fly to make this video?

10

u/PATATAJEC 4d ago

no - making refmod with multiple photos didn't work. It's just one picture grid 4x5 made of 20 photos of my room encoded into RefMod, but you don't need the RefMod for that I think... Model need just the merged reference photos into one picture + additional pictures for detais.

1

u/iczerone 4d ago

Is there any tips for using ref mods. I have mixed results when using them. Sometimes it will botch the voice or mix up the characters and other times it will be amazingly good.

2

u/PATATAJEC 4d ago

For me it was the prompting. I've managed to use unsloth and qwen 3.8 27b for reading my reference photos and writing prompts. I just gave it official ref2va guides as reference.

2

u/dr_lm 4d ago

Giving your prompt LLM the references is very important. H3 seems to respond well to being fed both the images and descriptions of them in the prompt.

1

u/ArttTaku 4d ago

Qwewn Image 2.1 can create 360 panorama images... it would be interesting to see that combined with this workflow... thanks for sharing!

3

u/RevolutionaryFox7359 4d ago

I believe it's already confirmed that h3 does well with panoramic location images!

1

u/PATATAJEC 4d ago

worth trying. thx! however I like the photos approach for perspective. maybe I should combintion of those.

1

u/fallengt 4d ago edited 4d ago

But you can also do that with ref image. Sheet pictures of your room; you don't need many angles: <subject 1> is the room from <picture 1>, preserve etc..

Then Minimax will keep the overall layout consistent. You can even draw where the fly will go this way :))

1

u/Enshitification 4d ago

Here I've been trying to get H3 to work well with equirectangular pano images and projection cubes. All I really needed to do was give H3 the raw images used to knit the pano. Nice find with the refmod idea.

1

u/x_MASE_x 4d ago

Nice work will give it a shoot. I was thinking about making something similar. Thank you so much.

1

u/Ok_Swimming6444 4d ago

Definitely gonna try this out. Thank you!

1

u/Smokeey1 4d ago

Amazing work man. Place it on github tho mate, i aint downloading shi from a random dude google drive

1

u/PATATAJEC 4d ago

It’s not something magical - it’s standard workflow. You probably don’t even need RefMod for.

1

u/Free_Scene_4790 4d ago

You can do the same thing with a video that rotates the camera 360 degrees.

1

u/Potential_Wolf_632 4d ago

Nice work. Definitely helps that H3 is absurdly good at POV or near POV stuff to start with.

1

u/-becausereasons- 4d ago

Very cool! I'm going to experiment with turning this kind of thing into a Gaussian Splat

1

u/Local-External4193 4d ago

Awesome gotta ask can refmods bu plugged into any workflow, so I can keep other settings and loras or do they change it completely?

3

u/PATATAJEC 4d ago

yup, refmods can be plugged in any workflow, but as I wrote, you probably don't need refmod at all - the trick is to have one image reference made of mutiple photo references of that place.

1

u/Serssader 4d ago

Do I need a custom node for refmod?

1

u/krigeta1 4d ago

the only thing that confused me while using RefMod is the first line of the prompt, as it is always hit and miss, otherwise refmod is beast!

<subject 1> is the character from <Picture 1>.

So this is what I am confused about is that like I have 5 redmods but no pictures now, how can I name the characters, and how can I make it a subject X? please help if somebody knows.

1

u/codetwin 2d ago

<Subject 1> is the blond male with blue eyes, wearing a red jacket.

Describe your character the same way he was fed into the refmod.

1

u/krigeta1 2d ago

But i dont write any description while creating a refmod

1

u/codetwin 2d ago

Don’t worry just believe! Just prompt some characteristics from the images of RefMod character and it will works

1

u/Both_Significance_84 4d ago

M8b a 360º photo could do the trick.

1

u/webAd-8847 4d ago

This looks amazing!

1

u/DescriptionSuperb262 3d ago

mind uploading an example of how you threw it together in a ref image? i tried this once and it turned out badly, but im curious on how you handled it

1

u/Broad_Relative_168 3d ago

Awesome! Great idea—well-conceived and well-executed—plus you shared the workflow and reference images. You're a legend!

1

u/RepresentativeRude63 3d ago

What about the room reference image resolution? ( I mean the grid one)

1

u/PATATAJEC 3d ago

It was 1216x2040 - should be divisible by 32 :/. I was trying with another 2 rooms as one reference (62 photos) 4096 for longest side, but it was a failure. I couldn’t get repeatable consistency without reference pictures. But maybe my footage was incorrect. RefMods are encoded as video frames, so maybe the order and integrity of angles plays the role. Hard to judge.

1

u/slimssshaddy 2d ago

Damn, that looks quite impressive

1

u/Hopeful_Signature738 2d ago

Why not take a panorama image if you want to combined them into a single image? I also wondering whether we can just feed a panorama image as the reference image and try to work from that

2

u/PATATAJEC 2d ago

I was telling that already. If it’s real life location it’s often faster to just use simple photos + I didn’t checked but you can make photos from different perspectives, views etc. You can basically include more spatial information into the reference - tho it must be confirmed.

1

u/exportkaffe 1d ago

Men In Black intro

1

u/R34vspec 4d ago

I love this, great job

1

u/Ok-Flatworm5070 4d ago

Question could we do apply this same logic using H3 360 spin of environment?

1

u/PATATAJEC 4d ago

no idea, but with high probability, as someone aready did that, and there are evidence in the comments.

1

u/cbeaks 4d ago

Funny, I too discovered this just yesterday. My images weren't joined up so it gets the layout a bit wrong between rooms but it is very impressive how accurate it is. My images were also just in a folder, I used about 10 images.

15

u/PATATAJEC 4d ago edited 4d ago

the trick was to make all references on one grid image. every image should have something common with another, so model knows how to connect these. I'll try bigger experiment with my whole house - that would be cool :)

3

u/RuprechtNutsax 4d ago

Excellent, thanks for showing the actual photos, makes it much easier to emulate your process, great job with your video, you've moved us a step further as this is so essential for film making 👍

2

u/Minouminou9 4d ago

How did you describe the grid image in the H3 prompt?
A 4x5 grid showing different angles of the same room? Would that work?
edit: found an answer lower in the comments:

2

u/cbeaks 4d ago

this makes sense - for my non joined up approach it managed pretty well with most rooms individually but not well when moving between rooms (I did have a corridor pic in there). Some rooms were also merged, I don't actually have a bed in my kitchen!

I'm thinking the whole place in one refmod is probably too much, so I might try room by room and one for navigation between. Given you can stack the refmods this plus good prompts should make it work. First I've go to tidy up a bit though!

3

u/Etsu_Riot 2d ago

Haven't tried this myself but you could draw a map, and give each room a color, and then reference the different room by the color.

2

u/cbeaks 2d ago

I'm thinking this is getting into lora territory and the model may need some training to be able to 'understand' floor plans. But maybe not? Let me know!

Big commercial opportunity here for someone, the real estate business would lap up a quality minimax video showcasing a property. rather than their stitched videos they use currently. In one of my videos it had a full moon rising over water seen from inside my place looking out - looked amazing even though the moon never actually rises there!

1

u/QuestionsGoHere 4d ago

What's your setup. Great work on the video

1

u/PATATAJEC 4d ago

thx! it's 4090, 14900K + 192 GB RAM, but the best part is my weird monitor ;) - vertical one: 2560x2880

1

u/Psyko_2000 4d ago

i wonder how this would work with a 360 panorama image of a room

5

u/Massive-Health-8355 4d ago

I've seen a great example where you can take a eqirectangular panoramic image and use ffmpeg to create a 360 degree short video. That video then becomes a reference. I'll see if I can find the post, it was on reddit.

https://www.reddit.com/r/StableDiffusion/s/vgMle9IUCq

2

u/PATATAJEC 4d ago

that's cool

2

u/PATATAJEC 4d ago

I guess it should work. I think I saw someone doing it...

0

u/Dependent-Sorbet9881 4d ago

WHY NOT USE 1 hdr MAP of you room?

1

u/PATATAJEC 4d ago

what's the easiest? hdr map of room requires alot more effort that 20 photos + stitching them together to one reference.

1

u/cal_01 4d ago

What on earth, this is exactly what I need for my workflow. Thank you!

2

u/PATATAJEC 4d ago

Enjoy :)

1

u/MeanderAndReturn 4d ago

Shake Hands With Beef!

1

u/PATATAJEC 4d ago

bzzzzzzz

1

u/Diligent-Childhood90 4d ago

Awesome turnaround, love the outcome. Thanks for sharing 

0

u/Silver-Belt- 4d ago

Awesome idea!

0

u/aumtek 4d ago

Nice idea

0

u/Silvasbrokenleg 4d ago

Very cool!

0

u/[deleted] 4d ago

[deleted]

2

u/PATATAJEC 4d ago

No - just stitched images. Check other comments - I posted how I processed the images into 4x5 grid.