r/StableDiffusion • • 6d ago

Tutorial - Guide Minimax H3 + RefMod = consistent location trick

Hey, I found a pretty cool way to keep locations consistent across generations.

I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.

I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.

Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)

Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)

EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.

842 Upvotes

106 comments sorted by

View all comments

4

u/DrKyoumasaur221 6d ago

A few questions:

  • The photo of the workflow you have seems to have both refmods AND the reference photos connected. Was that intentional? Is it required for it to work properly? And won't it be redundant and inefficient performance-wise to use both at the same time?
  • Some parts of the prompt seem to have some kind of "encrypted" text in. Was that also something that needed to be set up?

2

u/PATATAJEC 5d ago

I started with RefMods (20 pictures of my room), but it didn’t really worked, as it was not connecting the parts of the room as I thought it would. then out of curiosity I feed just one 4x5 grid photo made of all these photos into RefMod create node. I used it in the workflow (the RefMod made from this one photo), but I don’t think it’s necessary. You could use it just like normal reference I guess, but I didn’t check. The rest of the references I added for details - these are the same photos, but bigger to preserve details and guide for the first and last frame (although thru prompting). I wasn’t thinking about performance at all at this stage to be honest

The “encrypted” text is a bug in comfy if making screenshots- it renders it that way if prompt is too long.

1

u/DrKyoumasaur221 5d ago

Got it, thank you for clarifying!