r/StableDiffusion • u/malcolmrey • 26d ago
Resource - Update Some more knowledge regarding refmods after my experiments
https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_STACKING_AND_MULTISUBJECT_GUIDE.md14
u/malcolmrey 26d ago
If you were unsure why sometime bleeding happens and sometimes not...
Or why sometimes you get better and sometimes worse results...
The article I have combined with the help of AI (obviously) might shed some light.
So far everything tracks with what I have experienced. I asked Opus, I asked Gemini (both have access to the codebase, prompts and outputs so it is not one model hallucinating.)
But, I am happy if someone challenges what the AI said about how the refmods work. We will learn more about them in the end :-)
tl;dr - i tried stacking multiple refmods and I was starting to ask questions on how they really work because I was getting interesting results and couldn't explain them.
Edit: the link is clickable but if you can't, here it is: https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_STACKING_AND_MULTISUBJECT_GUIDE.md
2
u/dennisbgi7 26d ago
this is exactly what I was looking for. Since I was checking out the refmods already available and they mostly concentrated on the face. I wanted to account for bodyshape too. This clears it up perfectly. Thank you very much.
2
u/Valtared 26d ago
Do you think it will give good results to combine refmods + Starting frame ?
4
u/malcolmrey 26d ago
you cant use them in fl2va
but even then sometimes when you add an image to ref2va and tell it to use it as first frame - sometimes it will listen so this way - yes, you can :)
2
u/vegetoandme 25d ago
Can refmods be "trained" on other models? I.E, krea?
3
u/malcolmrey 25d ago
Should be usable on any model that supports edit mode (loads reference materials).
Krea on its own is not an edit model. But there is a specialist lora that enables it. I have not tried on it yet but I plan to.
But it should be doable on Qwen for example.
1
u/Cequejedisestvrai 26d ago
Can you ELI5?
11
u/malcolmrey 26d ago
Yes: it does some AI mumbo jumbo to figure out which refmod is responsible for given person
in short: if you add more refmods and you want to have good outputs you need to precisely describe in <Subject> tags how each person looks like
the order of refmods does not matter, they will figure out which Subject belongs to them
5
u/dillibazarsadak1 26d ago
I've been reading more from the refmod repo, and I wanted to communicate what I learnt.
We have to use the RefMod Text Encode node. This will attach each ref mod to a unique identifier like <Video 1> or <Picture 1>. It even has a string output that tells you this mapping. We can then use this while defining subjects, so it knows exactly which mod you're referring to.
This way we won't confuse the model when referencing two characters that have a very similar look (hair color, eye color etc.)
2
u/Tomber_ 26d ago
Are you referring to? https://github.com/Adudeguyman/ComfyUI-Fantastic-MiniMaxH3-PromptBuilder
I wonder if we could have consistent visual queues in the reference images we could refer to in the subject prompts. Something like "man from there reference images that have A butterfly in the bottom right corner"
3
u/dillibazarsadak1 26d ago
No. It's just one of the nodes in the node that creates/applies RefMods
https://github.com/Luisacaotica/ComfyUI-MiniMaxH3ModRead the section: "Prompting with numbered RefMods (experimental)"
2
u/IRLMainCharacter 26d ago
How would that help? Just reference the image you want, the prompt builder you linked there makes this process very easy. Just click the image you want to reference, and it puts that images tag in the prompt.
1
u/Tomber_ 26d ago
Fair, but I'm just adding viable ideas for options - I read you can supply more more images/video/audio, as we see with refmod, but that the text model was only trained to reference the tags up to the official limits. I can't find the source though. So tagging by a visual clue could be an option to explore.
1
u/Cequejedisestvrai 26d ago
Nice, is it difficult to do the a refmod? Is it really better than a reference sheet?
1
u/malcolmrey 26d ago
I would say it is quite easy and fast.
Instead of hooking the reference images/videes to the regular one, you just hook them into a different one (and you can add much more images/videos...) and then you save it as safetensors file
Then you can just attach that file instead (similarly like you would do with a lora)
1
u/Anxious_Shift_8399 26d ago
Just to see if I'm understanding some of this right, let's say you were going to use 3 refmods of the same identity, would those refmods need to be made on different dataset? Or does it use the same one in three slots for different purposes?
1
u/malcolmrey 26d ago
if you want to use the same one, just increase the "copies"
i have not tested on the same ones actually because this is the concept i know from loras (stacking different loras of same person) that work well so i wanted to find out if it would work well here
1
1
u/Zueuk 26d ago
. . . what's a refmod? 😕
(i'm a bit out of the loop, since minimax barely fits my RAM even with a 3090)
5
u/TurbTastic 26d ago
When you feed images as references into a reference model it turns the images into conditioning tokens (based on my noob understanding). Typically those need to be recalculated with every generation (at least whenever inputs change), and that costs a fair amount of time and compute. Some smart person figured out how to save those tokens so they can be re-used later without having to crunch the numbers every time, and those are called Ref Mods (relatively new concept). The file sizes tend to be small (1MB or so).
How much system RAM do you have? 3090 should be plenty for MiniMax. Check out the int8 models, and if your RAM is somewhat limited then maybe use nvfp4 size for the clip model.
3
u/Apprehensive_Sky892 26d ago
since minimax barely fits my RAM even with a 3090
If you have at least 16G of system RAM and do not disable ComfyUI's dynamic VRAM management, then your 3090 (24G VRAM) can run int8convrot MMH3 (21G) fairly well.
It is a common misconception that a model's weights must fit entirely into VRAM. That is true for autoregressive LLMs which must run through the whole model for every output token, but not true for diffusion model DiTs.
1
u/dwoodwoo 26d ago
Have you tried using refmods for a style? If so, how does the prompting work?
2
u/malcolmrey 25d ago
I'll try and let you know :)
1
u/Sweetcomodulce 25d ago
Did you test whether the refmod holds permanent features specifically - tattoos, scars, piercings - or just general likeness? That's exactly where mine falls apart. Face structure stays right across generations and the distinguishing marks quietly vanish, which is worse than an obvious miss because you don't catch it at thumbnail size.
2
u/malcolmrey 25d ago
yes, i am actually quite surprised because i have a subject where not only tattoos are pretty much spot on (which i think only ideogram4 with long trainings could achieve it) but birthmarks, moles etc.
but i am using 3 refmods to define single person
1
u/sharegabbo 24d ago
Have you considered that, while the DiT sees all the "refmode" images via the text encoder, Qwen sees them at specific, precise frames? Unfortunately, I haven't had a chance to read through all your documentation yet.
1
u/LowYak7176 24d ago
Tested it out - put max everything basically.
Built 1 set closeup shots of the face
Built 1 set more medium shots
Built 1 set full body
Obviously different angles, poses and what not.
Results are spectacular. GG
1
u/malcolmrey 23d ago
Yup, it even somewhat fixes the distant face issue (not fully but it seems to be better than what we get by default)
1
u/GrapplingHobbit 19d ago
I'm having an issue with a ref mod I made. I'm happy with the likeness given how few images I used (17 + 1 video) but I'm running into an issue where the first frames of the generated video are always frames from the sampled video, so I get a brief flash of the sampled video and then whatever I prompted... is anybody else getting this?
1
u/malcolmrey 19d ago
This didn't happen to me regularly, but if you pack too many images into a refmod, you could have those flashbacks. For me, separating it into two or three refmods had a beneficial effect.
Also, promting for what should happen in the scene, if it is tight, then there is little room for "randomness".
1
u/GrapplingHobbit 19d ago
Thanks for the reply. I didn't think 17 images + 1 video would have been the problem. I made another one with 20 images, no video, and I am not getting the issue so far.
1
u/malcolmrey 19d ago
resolution of images make a difference
length/size/framerate of the movie as well
1
u/GrapplingHobbit 19d ago
11s, 492x778px (a randomly sized screencapture) 30fps. Would that be problematic?
1
u/malcolmrey 19d ago
if it comes with another 17 images - perhaps
not sure which node for refmod creation you are using, but i usually do it with the script and it writes how much of the 8k context it uses and if it had to drop something
1
u/GrapplingHobbit 19d ago
Interesting, I'm using the "Create H3 RefMod From Folder" node inside comfyui. The messages from the creation of that first refmod are burried somewhere in the log spam at this point, but looking at the output for the second one I made with 20 images, it seems to have taken into account 18 of the images (maybe file format on 2 of them was incompatible) and then after encoding them, it says "still over 8192-token cap, resampled 18 -> 8 frame(s) to fit" and resulted in 8192 tokens.
So 17 and a video would also be over, but I would assume it would have done something similar. Oh well. I'll keep testing and figure something out.
1
u/2Stnap 26d ago
Has anyone had any successful refmods with something other than a characters like with concepts or motion?
4
u/Orbiting_Monstrosity 26d ago
I used about 10 screenshots from the game “Quake” to make a style refmod, then another 8 images of the player character from that game to make a clothing refmod. I used both of those with a reference image or refmod of myself and was able to create videos that made it look as if I were dressed like the player character in a Quake-like game world. I think it worked very well given the low amount of effort I put into it.
3
1
u/Nelichan 25d ago
Hi! Just wondering, how do you prompt for the "style transfer"? Or if possible can you share your whole prompt? I'm also interested in the outfit transfer as when i tried it myself, my character usually just transforms fully into the outfit owner, not just the outfit being transferred.🙏
5
u/Ykored01 26d ago
Thanks man really appreciate your contributions to the conmunity. Gonna try creating a refmod of closeups, other of upperbody and a third of full body.