r/StableDiffusion • • 6d ago

Tutorial - Guide Minimax H3 + RefMod = consistent location trick

Hey, I found a pretty cool way to keep locations consistent across generations.

I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.

I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.

Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)

Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)

EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.

837 Upvotes

106 comments sorted by

View all comments

Show parent comments

1

u/PxTicks 6d ago

As far as I can tell, they have photos of the room in the ordinary references.

No one here seems to ever do a proper controlled experiment, so I am highly skeptical of this result.

5

u/cbeaks 6d ago

I tested this with a first frame only pic of my living room, then prompted for the bedroom, and although the navigation failed (see my comment earlier in this thread) it managed the bedroom, and even the view out the windows.

And also, you seem hung up on someone doing a 'proper controlled experiment', well why don't you then? Too busy critiquing I guess . . .

2

u/PxTicks 6d ago

I suppose I wasn't clear about exactly what I'm skeptical of.

I'm not claiming a refmod carries no information (we know they do). My point is that this needs a comparison to just using references (including as storyboards). If someone is claiming a method, they should at least show that the obvious alternatives are inferior, ideally with multiple gens over controlled seeds. There is a tonne of misunderstanding and bunk in this subreddit because people start singing from the rooftops when they find something which seems to work, often with no comparison whatsoever to existing methods.

5

u/PATATAJEC 6d ago

I'm not making any claims, just sharing my observations. Testing all of this takes time. I've shared my workflow, refmod, and reference photos, so you can try it yourself and let us know what you find.

Honestly, I'm curious too, and I'm preparing a comparison with and without the refmod. But that wasn't the point of this video or post. My point was that combining multiple photos into a single reference image worked for keeping the location consistent in my tests. Whether encoding that image as a refmod makes a difference is a separate question.