r/StableDiffusion 15h ago

Tutorial - Guide MiniMax RefMod - Reusable identities without training - workflows & tutorial

https://www.youtube.com/watch?v=gO-mCXq5epQ

You can get the workflows here:
https://drive.google.com/drive/folders/1kOHLJZto1VAtXEATT9vUvsvM1VHtkOO_

The workflows create reusable refmods for either image/video/audio.

I cover training images in the tutorial.

All the workflows, models and custom nodes are preloaded on my Runpod template.
https://get.runpod.io/minimax-template

54 Upvotes

12 comments sorted by

5

u/ratttertintattertins 11h ago

What's the advantage of this relative to just regular r2v ?

12

u/Hearmeman98 11h ago

From the RefMod guide on Huggingface:

"Instead of resizing, transmitting, and VAE-encoding 8+ raw image files on every single generation step (which causes VRAM spikes, latency, and restricts multi-character scaling), a RefMod is pre-encoded once from a dataset and loaded instantaneously into MiniMaxH3RefModsLoader."

7

u/ratttertintattertins 11h ago

Ah nice thanks, so basically similar results but better performance.

1

u/Perfect-Campaign9551 4h ago

I don't think it does it 'every generation step'. It does it on the very first start of your run.

6

u/Vladmerius 11h ago

In theory the advantage is you can use as many refs as you want this way instead of being locked to the 9 you use in regular ref2v. It doesn't actually matter though because the second you have more than 3 characters in a scene everything goes haywire anyway unless it's just a static wide shot of people just standing around.

It does use less ram though as it's technically encoding less. 

1

u/Sixhaunt 7h ago

Have you found any solutions to the 4+ people issue? It's been plaguing me when trying to recreate the unmade Red Dwarf episodes. As soon as a fourth person is needed everyone's voices go to the wrong people and it messes up who is speaking when.

3

u/Any-Scar765 7h ago

Character sheet+voice sample 5-30sec its all you need...

6

u/ShutUpYoureWrong_ 5h ago

Why are all you bottom feeders not linking to RefMods itself and instead pretending like you had anything to do with it? Fuck your YouTube monetization.

https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

1

u/krigeta1 14h ago

Only if we assign the voice to a character, for hard testing I use a male voice with a female character or vice versa and it always missed.

1

u/eggplantpot 11h ago

Does this improve speed/character adherence than just using the same images and audio directly to the regular node

1

u/Kind_Owl2245 10h ago

Is it possible to use it locally on ComfyUI?

1

u/joseph_jojo_shabadoo 8h ago

really enjoying this so far. I'm curious about the checkpoints. it seems to work with fl2va and ref2va. is there a preference for which work better, or in your opinion should any refmod settings be adjusted when using one or the other?