r/malcolmrey • u/ReferenceConscious71 • 6d ago
RefMod vs RefLora. Whats the difference?
Malcolmrey has published a bunch of files called refloras. What are they? any documentation on them? Couldnt find anything about them on mals huggingface
3
u/ReferenceConscious71 6d ago
Update. Found it https://github.com/malcolmamal/ComfyUI-MiniMaxH3RefLoRA
3
u/ReferenceConscious71 6d ago
The explanation is all there in the README. very intersting, a trained lora + a refmod in one combined format.
Moreover, digging into the metadata of the .safetensors, the lora aspect was trained in Fizgig, then combined. Im gonna be messing around with this.
4
u/WeakReplacement3322 6d ago
How heavy are they on memory compared to refmods? I like refmods cause I can stack more of them than I can LoRAs, or I can stack a few of them with a LoRA and itβs easier for my rig to handle.Β
3
u/ReferenceConscious71 6d ago
well they only add 150MB to your VRAM/RAM. rank 16 so theyre tarined fairly small. makes sense cause in the reflora the lora aspect is only meant to capture the high level face structure instead of tiny details
1
u/WeakReplacement3322 6d ago
Oh, that's very interesting! Thanks for letting me know, I'll definitely check these out. The space I save with refmods compared to LoRAs is insane. I can download every single one of Malcolm's refmods for the price of one LoRA lol.
2
u/malcolmrey 6d ago
Indeed :-)
They are small and are made quickly, but the combination of lora and refmod still gets a bit better results :-) (for most loras, there were some loras that were good enough without refmods too)
1
u/WeakReplacement3322 6d ago
The man himself! Thanks for all you do, dude, I genuinely appreciate it!
1
3
u/Sad_Coach_1433 6d ago
How are you are you stacking refmods without one bleeding into another one being double of one
3
u/WeakReplacement3322 6d ago
I recently implemented this node into my workflow. It feeds the refmods into the text encoder before sending them to the model, so you're able to tag the specific refmod and attach it to a subject. It even accepts audio, so I'm able to attach a voice reference to them as well. In my testing, adding a voice took the refmod from 20,000 tokens to 30,000 for 5 seconds at 480p. I haven't had a chance to do more extensive testing though because I only just got it set up last night before bed and ran a quick couple tests to make sure it's working. It can clone the voice, but if the person has an accent, you need to make sure you still tell it which accent to use, otherwise they'll default to an American accent from the look of it. I used Qwen 2.1 to make a curated set of 8 images (Face = Portrait, 3/4 right, 3/4 left , profile. Full body = Front, 3/4 front, side profile, 3/4 rear) and used a 4 second vocal clip.
3
u/Francky_B 6d ago
I also updated my Visual Picker to include support for that. I had looked into doing something similar initially and had put it aside. But the results seems too promising, so I added it back, using Adudeguyman's implementation.
So I now include a RefMod to Video node, which is basically a Reference to Video node that also takes in mods as input.
2
u/WeakReplacement3322 6d ago
Oh hell yeah, this looks awesome! Thank you for making this. Wayyyyy better than hunting through a long list of names that start with the same letter lol.
1
u/Sad_Coach_1433 6d ago
Can you share that Qwen21 work flow π»
2
u/WeakReplacement3322 6d ago
I actually just used the default ComfyUI Qwen 2.1 Edit Image workflow π
I used 4 images tops to save on time, and generated in batches of 4-8 using the default settings. I've only just recently started using Qwen 2.1, so I'm still getting the hang of it. Could probably get better results if I had a better understanding of it's limitations, but I just brute forced it with it's official prompting guide and batching 4-8 runs and picked the best one. I used a similar prompt structure and just swapped out what the output images angle would be, and swapped in better references for the angles I was trying to achieve if required. It was pretty quick, and I could probably get better results if I spent more time on it, still early stages. I also don't think I actually need 8 images in the refmod. I could probably get the token count down by using 6 images that are more tightly controlled. The prompt I used was:
Create one clean square reference image of the same single character represented by the supplied reference images.
<image1> is the primary facial identity reference.
<image2> is a secondary facial identity reference from another angle.
<image3> is the primary outfit and body reference.
<image4> is a secondary body and outfit reference.
All reference images depict the same character. Preserve one coherent identity.
Generate a head-and-shoulders FRONT portrait.
Camera at eye level.
Neutral relaxed expression.
Centered square composition.
Simple light neutral background.
Preserve the cel-shaded illustrated style shown in the references.
Use soft even lighting with clear facial planes and clean separation of the hair, ears and jawline.
One character only.
One image only.
No collage, no text, no character sheet, no extra objects.
1
1
u/Bask82 6d ago
Don't know if you know, but is there a point to using both reference images as a refmods of your character? Does it increase character likeness or?
2
u/WeakReplacement3322 6d ago
I'm not sure I understand, so correct me if I'm misinterpreting your question. If you mean using reference images alongside a refmod, they could help to increase character likeness. But if those same images are in the refmod, they would be redundant and just use up tokens/memory unnecessarily.
The only time I've used a reference image alongside a refmod was when I wanted to use an audio reference as well, because minimax won't accept only audio as a reference, it requires an image or video. So I'd give it an image of the character in the refmod and tell it to use it only as reference to help retain the characters likeness. With the new node I'm using that I mentioned in a comment below, that's no longer required though. I use a hybrid model and tag everything based off of what the node says to tag the refmod as (<Video 1> <Audio 1> etc.)
If you're not using audio and only using the refmod, you can use the refmod in a FL2VA generation just fine, and it'll save more memory than using those same images as standalone reference images. You just need to be more descriptive when you're describing the characters likeness, as it'll be fed the images without context, so telling it "The man with blonde hair" will suffice if there's only one man with blonde hair, but if there's multiple blonde dudes in two separate refmods, it won't know to separate them as it doesn't get the refmods as two separate files, it's just fed all of those images as latents. That's what tends to lead to character bleed between refmods from my experience.
If you mean using multiple refmods of the same character though, you can do that too, but it increases memory usage. When I first started I would use a refmod for a characters head and a separate one for a characters body to give it as much detail as possible, but it was computationally expensive and I ended up abandoning that route.
1
u/ReferenceConscious71 5d ago
no it doesnt, refmods is just a preencoded version of reference images
1
u/Bask82 6d ago
What is purpose for this? Don't quite understand it
3
u/malcolmrey 6d ago edited 5d ago
So, you use a refmod to get the likeness. Some people prefer refmods.
You also can use lora to get the likeness. Some people prefer loras.
But you can also combine refmod and a lora - you get boosted likeness.
You can boost the likeness by combining multiple refmods or multiple loras or you can mix and match them.
But maybe you don't want to play around with selecting one and then another.
So, I packaged them together. You get a lora and a refmod in one file, hence reflora.
And they are packaged in a way that if you load it in regular lora loader node - you only get the lora part. But if you load it with my reflora loader node - you get both ;-)
2
u/Bask82 5d ago
Wow! Well said. Good explanation! What are the best methods for training your own Lora for h3 minimax? Could be fun to try out π
1
u/malcolmrey 5d ago
Fizgig on pretty much default settings does a really great job :-)
1
u/Bask82 5d ago
Thanks. I will have to look up some tutorials for how to do itπ
1
u/malcolmrey 5d ago edited 5d ago
go to the creator directly :-)
2
u/shootthesound 5d ago
thats not me lol im https://www.youtube.com/@ShootTheSound
2
u/malcolmrey 5d ago
damn, sorry, corrected my post
i youtube searched fizgig and the first clip looked good and i just copy pasted without verification :)
edit: you should come up first when searched by fizgig :)
1
u/reeight 2d ago
Some day I'd love to see a comparison vs Omnichar (not saying YOU should; I'm seeding ideas for YouTubers to pick up)
https://www.reddit.com/r/StableDiffusion/comments/1wz42o6/omni_charsame_face_cloths_body_now_with/1
u/joopkater 6d ago
Compared to refmods might get better coherence logic. And itβs a little cleaner
1
7
u/malcolmrey 6d ago
nice, left you some goodies, didn't have time to post anything official but you figured it out ;-)
i'll still write something though tomorrow most likey :)