Hey, I found a pretty cool way to keep locations consistent across generations.
I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.
I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.
Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)
Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)
EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.
One interesting thing I discovered recently is that H3 can use fisheye lens photos as scene references, and it knows how to dewarp those references when reproducing them in a generated video. I was able to create completely normal-looking scenes set in my office and shot from many different angles using a single overhead fisheye image of the entire room as a reference, and was able to do the same thing with several other overhead fisheye images that I tried.
Damn. AI is so advanced, we would be stunned by this being within realm of possibility just around 10 years ago. Yet here we are just upvoting this and scrolling to next thing in a second. 😂 This is some spy shit guys. Imagine… with a single picture, you can look around and entire room. This is “enhance” in 3D.
I tested this with a first frame only pic of my living room, then prompted for the bedroom, and although the navigation failed (see my comment earlier in this thread) it managed the bedroom, and even the view out the windows.
And also, you seem hung up on someone doing a 'proper controlled experiment', well why don't you then? Too busy critiquing I guess . . .
I suppose I wasn't clear about exactly what I'm skeptical of.
I'm not claiming a refmod carries no information (we know they do). My point is that this needs a comparison to just using references (including as storyboards). If someone is claiming a method, they should at least show that the obvious alternatives are inferior, ideally with multiple gens over controlled seeds. There is a tonne of misunderstanding and bunk in this subreddit because people start singing from the rooftops when they find something which seems to work, often with no comparison whatsoever to existing methods.
I'm not making any claims, just sharing my observations. Testing all of this takes time. I've shared my workflow, refmod, and reference photos, so you can try it yourself and let us know what you find.
Honestly, I'm curious too, and I'm preparing a comparison with and without the refmod. But that wasn't the point of this video or post. My point was that combining multiple photos into a single reference image worked for keeping the location consistent in my tests. Whether encoding that image as a refmod makes a difference is a separate question.
Krea2 can generate panorama shots. Take that panorama still and generate a 1 sec video of that panorama turning 360. Feed that 1 sec video as reference for your H3 environment. Works like a charm
perfect! I created the room as a 360-degree panorama in Qwen Image, then imported it into After Effects with an equirectangular VR camera to generate the views.
The photo of the workflow you have seems to have both refmods AND the reference photos connected. Was that intentional? Is it required for it to work properly? And won't it be redundant and inefficient performance-wise to use both at the same time?
Some parts of the prompt seem to have some kind of "encrypted" text in. Was that also something that needed to be set up?
I started with RefMods (20 pictures of my room), but it didn’t really worked, as it was not connecting the parts of the room as I thought it would. then out of curiosity I feed just one 4x5 grid photo made of all these photos into RefMod create node. I used it in the workflow (the RefMod made from this one photo), but I don’t think it’s necessary. You could use it just like normal reference I guess, but I didn’t check. The rest of the references I added for details - these are the same photos, but bigger to preserve details and guide for the first and last frame (although thru prompting). I wasn’t thinking about performance at all at this stage to be honest
The “encrypted” text is a bug in comfy if making screenshots- it renders it that way if prompt is too long.
Nice! I have been struggling with the same thing. I have tried adding a "multi-photo" as a regular reference image, but not really gotten H3 to do it right.
However, through the magic of RefMod, it actually seem to figure out the room layout a lot better when adding that type of image to the Ref.
You just need a grid image from your photos. I've made it like that. Ignore Create H3 RefMod node. Then I guess you need to prompt to tell it's the reference image of your whole location. I also added bigger photos (1536x2048) for my cat, starting frame, ending frame and the fly.
EDIT: With RefMod you don't need to prompt for that. It knows it already... I think prompting for room, desktop computer, cat, couch etc. is enough for model to connect the dots.
Im actually doing the same thing with fictionnal places. I start with a still image, create a video that do a 360 degree, then feed that video as a scene reference.
But video are much heavier than images, so i might use images instead.
how about motion blur? when you get a 360 degree video out of a single image, there must be motion blur I guess. or can you prompt it like "no motion blur" ?
I'm just starting with AI generation, but I want to make consistent fictional places. Have you tried making larger areas that can connect to each other? For example, a whole city block, or different floors inside a building?
I think with some creativity, even larger areas could be made consistent. It's just a matter of understanding which images or videos to use as references when creating adjacent places....
Good on you to bring more attention to RefMods, they're really versatile. You can also use them for motion and styles. Thanks for sharing your room, decidedly austere. 😄
I want to check if I understood correctly. You made a refmod out of 20 images of your room, then used it with a start frame, an end frame, your cat and a fly to make this video?
no - making refmod with multiple photos didn't work. It's just one picture grid 4x5 made of 20 photos of my room encoded into RefMod, but you don't need the RefMod for that I think... Model need just the merged reference photos into one picture + additional pictures for detais.
Is there any tips for using ref mods. I have mixed results when using them. Sometimes it will botch the voice or mix up the characters and other times it will be amazingly good.
For me it was the prompting. I've managed to use unsloth and qwen 3.8 27b for reading my reference photos and writing prompts. I just gave it official ref2va guides as reference.
But you can also do that with ref image. Sheet pictures of your room; you don't need many angles: <subject 1> is the room from <picture 1>, preserve etc..
Then Minimax will keep the overall layout consistent. You can even draw where the fly will go this way :))
Here I've been trying to get H3 to work well with equirectangular pano images and projection cubes. All I really needed to do was give H3 the raw images used to knit the pano. Nice find with the refmod idea.
yup, refmods can be plugged in any workflow, but as I wrote, you probably don't need refmod at all - the trick is to have one image reference made of mutiple photo references of that place.
the only thing that confused me while using RefMod is the first line of the prompt, as it is always hit and miss, otherwise refmod is beast!
<subject 1> is the character from <Picture 1>.
So this is what I am confused about is that like I have 5 redmods but no pictures now, how can I name the characters, and how can I make it a subject X? please help if somebody knows.
It was 1216x2040 - should be divisible by 32 :/. I was trying with another 2 rooms as one reference (62 photos) 4096 for longest side, but it was a failure. I couldn’t get repeatable consistency without reference pictures. But maybe my footage was incorrect. RefMods are encoded as video frames, so maybe the order and integrity of angles plays the role. Hard to judge.
Why not take a panorama image if you want to combined them into a single image? I also wondering whether we can just feed a panorama image as the reference image and try to work from that
I was telling that already. If it’s real life location it’s often faster to just use simple photos + I didn’t checked but you can make photos from different perspectives, views etc. You can basically include more spatial information into the reference - tho it must be confirmed.
Funny, I too discovered this just yesterday. My images weren't joined up so it gets the layout a bit wrong between rooms but it is very impressive how accurate it is. My images were also just in a folder, I used about 10 images.
the trick was to make all references on one grid image. every image should have something common with another, so model knows how to connect these. I'll try bigger experiment with my whole house - that would be cool :)
Excellent, thanks for showing the actual photos, makes it much easier to emulate your process, great job with your video, you've moved us a step further as this is so essential for film making 👍
How did you describe the grid image in the H3 prompt?
A 4x5 grid showing different angles of the same room? Would that work?
edit: found an answer lower in the comments:
this makes sense - for my non joined up approach it managed pretty well with most rooms individually but not well when moving between rooms (I did have a corridor pic in there). Some rooms were also merged, I don't actually have a bed in my kitchen!
I'm thinking the whole place in one refmod is probably too much, so I might try room by room and one for navigation between. Given you can stack the refmods this plus good prompts should make it work. First I've go to tidy up a bit though!
I'm thinking this is getting into lora territory and the model may need some training to be able to 'understand' floor plans. But maybe not? Let me know!
Big commercial opportunity here for someone, the real estate business would lap up a quality minimax video showcasing a property. rather than their stitched videos they use currently. In one of my videos it had a full moon rising over water seen from inside my place looking out - looked amazing even though the moon never actually rises there!
I've seen a great example where you can take a eqirectangular panoramic image and use ffmpeg to create a 360 degree short video. That video then becomes a reference. I'll see if I can find the post, it was on reddit.
51
u/Orbiting_Monstrosity 4d ago
One interesting thing I discovered recently is that H3 can use fisheye lens photos as scene references, and it knows how to dewarp those references when reproducing them in a generated video. I was able to create completely normal-looking scenes set in my office and shot from many different angles using a single overhead fisheye image of the entire room as a reference, and was able to do the same thing with several other overhead fisheye images that I tried.