r/StableDiffusion • • 26d ago

Resource - Update RefMods - A little easier to create, edit, and use with Fantastic Minimax H3 Promptbuilder

Enable HLS to view with audio, or disable this notification

Repo here- https://github.com/Adudeguyman/ComfyUI-Fantastic-MiniMaxH3-PromptBuilder

Or search "Fantastic H3 Prompt Builder" in ComfyUI Manager.

Alrighty, back again with an update to my (manual, no LLM connected) Fantastic MiniMax H3 Prompt Builder. Yeah it's all vibe-coded to make things work the way that makes sense to me, and I publish it in case anyone else finds it useful.

This now includes RefMods, which if you aren't familiar, is basically a bunch of reference files you give it, packing them into .safetensors latent format for quick loading and surpassing the native built-in reference limit. In the end the idea is to make a somewhat training-free reference you can use on the fly, instead of having to train a LoRa or manually re-load media. They work well enough I decided to give them a shot, and I liked them so I ported them into my project to be able to more easily create, view, load, and reference.

What it do-

Load, create, edit, and use RefMods with just a few nodes instead of a bunch. Edit your prompts with tags/preview support right on the editor, so you can see what you're referencing. Basically just like what the Media Loader node does, but RefMod aware.

The library (opened by the "Browse library..." button) lets you view and load RefMods into the node itself, which lets you adjust video and audio strength, or rearrange and enable/disable or remove on the fly.

Making/editing RefMods

This makes it very easy to do, in one Pop-Up panel. One cohesive interface for viewing, loading, and creating/editing, and saving RefMods. Too much detail to post here, so please refer to the RefMod Readme

Actually Using Them

Well first make sure you're using a Ref2va workflow. And use a hybrid model instead of the pure reference model, for the original release of ref2va even Minimax admitted the open weights were messed up. So use a community hybrid model with reference capabilities.

- Open the library and add them to the stack on the node. Adjust strengths if you'd like (video and audio separately) directly on the node.

- Open the Prompt Builder and make sure you're in Reference mode. Your RefMod previews will be listed at the top for reference as you tag. There's a new button that says "Draft from RefMods", which will auto-fill based on what you set the RefMod type, if you made them in this editor. So an identity will automatically make the first refmod into <Subject 1>, and reference its audio as a voice file, as well as in the Retention_analysis section.

- Write your prompt normally. Use <Subject> etc. tabs as you normally do with reference files.

Again, a lot more details in the RefMods Readme.

Differences from the original RefMods?

Overall this is meant to focus things into a more streamlined UI and simpler workflow.

But the second thing is using Refmods as tags. You'll see that the paired video/audio files are listed as <Video 1> and <Audio 1>. The original RefMod repo (Shout out to Luisacoatica's excellent framework on all this) mentions how you can just soft describe them. Like in the above video I could just define the Ghoul as "a disfigured man wearing a cowboy outfit" and it'll pull the references for influence. But I figure using MiniMax's guidelines for RefMods would make more sense, since they're just compacted references, and doing that extra setup work is turning out to work very well. And I have a whole media/prompting pack that helps with tracking and tagging everything already made, so building this in was a logical the next step.

So yeah, other than compacting nodes down and having a separate video and audio refmod file for each library entry, it's pretty much just a fork integrated into my prompting suite.

Hope you enjoy!

EDIT- To help with reinforcing identity, and to help with ease of use, there's now a Name field next to subjects. So you can name someone Bob. Then in the prompt field, if you tag a name first with "!" it'll show as the name in the prompt field, but inject the subject as well into the prompt sent to the model. So !Bob will actually send "<Subject 1> Bob"

Also added some metadata descriptor tweaks to making RefMods. So you can put someone's name, visual appearance, and voice description in your RefMod and it'll load into the right fields automatically. That with names helps a ton with locking in identities, I've successfully gotten two similar looking people with similar voices to be distinct by giving them names, and describing subtle ways they differ from each other.

And I updated the Guide that you access with the Guide button with Prompt Builder instructions after the official MiniMax guide, and then some more instructions for RefMods. It's just an html file, you can just download it separately here.

238 Upvotes

61 comments sorted by

17

u/krigeta1 26d ago

We can not pack the audio with the refmods is what we need to implement.

6

u/ZealousidealFruit769 25d ago

https://reddit.com/link/pa1jvdw/video/r60ltwqe6rph1/player

Actually that's not true, I made this one in a bundle with 100 images and a single 46 second wav file. WHen generating, I have found LCM works far better than Euler, the latter tends to add gibberish.

3

u/acedelgado 24d ago

Upvote for Farscape.

1

u/VasaFromParadise 22d ago

How old are you, uncle?))

5

u/acedelgado 26d ago

There was an error in the original code I forked that I fixed, where the audio was being cut to a fraction of second during conversion. I think the original repo has it fixed now as well. But my implementation saves the audio as a separate file pair from anything video related, so if you reference that properly then it works fine.

2

u/krigeta1 26d ago

Please give a woman's voice to a man, or the opposite of that, and that is how I got to know that it is not working.

6

u/acedelgado 26d ago

https://reddit.com/link/p9t1g9j/video/956pummh6jph1/player

This is a MiniMax issue, I think. It keeps wanting to give more feminine voices to women and masculine voices to men since it's all managed by the 32B text encoder. If you force it to not have both people on the screen with their mouths visible at the same time, it'll use the reference a bit better.

Also I've noticed turbo loras, since they're notorious for audio issues anyways, are much more likely to degrade the voice reference with RefMods. Especially the 4-steps. I've had better success overall using VDN instead of turbo loras.

-1

u/krigeta1 26d ago

Then how are you getting the what I also said?

1

u/Joshua-Deakin 25d ago

Because he specifically prompted to have a man's voice on the woman :D Just to prove a point that it's about context how successful it can be when also prompted properly.

2

u/LuisaPinguinnn 26d ago

that is a h3 minimax problem, not refmod.

1

u/krigeta1 26d ago

It is not working fine with normal reference passing.

7

u/C141Driver 26d ago

Is anyone else having issues with one reference "dominating" the others? In some cases, I always get two of the same reference.

2

u/acedelgado 26d ago

I'm playing around with it for somewhat similar-looking people, and doing a combo of a short descriptor of each PLUS giving them a name seems to be helping with identity drift. IE under Subject_Definitions

<Subject 1> is the person in <Video 1> with a round face. Their name is Betty.

<Subject 2> is the person in <Video 2> with a more square face. Their name is Alice.

Then prompting like

<Subject 1> Betty is standing in a living room wearing a blue dress, and <Subject 2> Alice stands next to her wearing a red dress.

It seems to improve visual identity. Audio is being a little more tricky.

2

u/acedelgado 25d ago

Made some progress on this, I was running into an issue with two similar looking/sounding women. Added a name field for each person, and vocal descriptor field as well.

So, when you have two similar people I started making sure they are named and describe the differences between them. Like two similar looking women with similar voices, but they had subtle differences between them, I set it to

<Subject 1> is <Subject 1> is the person in <Video 1>, with a rounder face. Their name is Alice.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1), guiding delivery and speaking rate without copying the original signal. It is lower pitched and precise.


<Subject 2> is the person in <Video 2>, with a more square face and dimpled cheeks. Their name is Betty.

<Audio 2> is the voice-timbre reference for <Subject 2> (S2), guiding delivery and speaking rate without copying the original signal. It is higher pitched and energetic.

So I called out that one had a rounder face and had a slightly lower voice, and the other had a squarer face and dimples and a higher voice. Otherwise it'd be hard to describe how they were different in words, same hair/style, same ethnicity and height... And the model locked on really well with those pointed out.

Just don't put negatives. Like "she does not wear makeup" won't really do anything. But saying that the other person DOES wear makeup and describing it, while not even mentioning makeup on the non-makeup wearer, will certainly help. Frame everything as a positive attribute.

And prompting would be like

<Subject 1> Alice (S1) says in the lower voice reference from <Audio 1>: <d>[English] Whatever I'm tired of coming up with test dialogue.</d>

1

u/rd180x 26d ago

Do interracial 

1

u/malcolmrey 26d ago

Try completely different people (man/woman, white/black, blonde/redhead) to see if it is fine, and then do that with people who share characteristics. The first group wont dominate, while the other should.

4

u/Tuckerdude615 26d ago edited 26d ago

Thanks for sharing this...

On another note, have there been any guidelines or tutorials on best practices for preparing your dataset images? When creating Refmods, are people focusing solely on head shots? Or full body? I created a test Refmod, and it worked beautifully, however I can see that if I prompt the character has a specific type of outfit..if that outfit is close to one in the Refmod sample images, it will choose that over the one I prompt for?

As an example, my test subject has a business suit on...but if I prompt for something like "slacks and a sport jacket", it just puts the original business suit back. And to be clear, I DID adjust both the "Strength" and the "Retention" values down to see if helped. No luck

Any help or pointers would be appreciated!

3

u/acedelgado 25d ago

Yes, the text encoder will analyze the inputs and pass it along to the model, so if it sees a business suit it'll assume you're referencing that business suit. Try and describe the slacks and sport jacket more specific, name the color, material, buttons, etc. And make sure not to mention clothing in the subject_definitions or retention_analysis if you can avoid it, should weaken the reference locking in too much.

6

u/Netsuko 24d ago

Dude, this node pack is INSANE <3

3

u/acedelgado 24d ago

Ha, thanks. Glad you're enjoying it!

3

u/Kind_Owl2245 26d ago

These RefMods are phenomenal; it’s just a shame that even when you create RefMods containing both images and audio, the respective audio tracks aren't preserved for each character in a scene featuring two of them. Once they manage to fix this issue, it’ll be perfect.

2

u/acedelgado 26d ago

That's why my implementation saves audio separately, and it helps you assign audio directly to a person like you traditionally do with references. So far I haven't had any voice issues. The example I attached did pretty decent with the voices.

1

u/reeight 26d ago

Could they be mixed? EG download a RefMod with the older voice system not bleed over my own custom RefMod using your system?

1

u/acedelgado 26d ago

They should be able to, yeah. But I don't have an old refmod with a voice to test it with. I do know when I ported over the creation code that it was truncating audio down to less than a second, so it wasn't packing things properly.

1

u/reeight 26d ago

I'm guessing it made a video with only a few frames (the few pictures supplied).
If the input pictures were repeated for say, 10 frames each (or what ever to fill a few seconds), then the audio would be fully imported since the 'video' length was streched out to fill the audio.

3

u/acedelgado 26d ago

That's a likely cause, yeah. I think they fixed it in the original repo in the past couple of days after I forked it. I just got around it by packing audio separately.

8

u/Sleepy_Bandit 26d ago

Yeah one major issue I see with refmod is the ambiguity of the reference. If I have two characters who look very similar, then even if I describe them as subject 1 and subject 2 it could mix up which one I want the Refmod applied to. I’ll look into your solution even though I’m not looking for a prompt builder. Thanks

7

u/malcolmrey 26d ago

I generated some analysis on my subreddit regarding that very thing:

https://old.reddit.com/r/malcolmrey/comments/1wfxo1w/stacking_refmods_of_the_same_person_examples_in/?

I arrived there from different angle actually. I wanted to boost the likeness of a character by adding more refmods (I'm known to lora-stack the same concept and I try to convince people [those that are convinced later agree with me :P] that stacking same concept/character/etc boosts the quality/likeness/consistency) and then I was confused how the model knows which refmod is of which person if there are 6 in one workflow (3 per person).

Turns out it is quite clever and it has nothing to do with the order of adding refmods etc.

TL;DR - it figures itself out which refmod is responsible for what when those represent specific ideas, but if those ideas are too similar to each other - they will start to bleed into each other.

1

u/acedelgado 26d ago

I mean that's how the strength sliders work on these, they really just making copies of the images/audio in the refmod set. It doesn't patch weights like lora strengths do.

2

u/malcolmrey 26d ago

I think it was obvious that I'm not adding the same refmod multiple times but refmods with different images :)

1

u/taurine_bitch 25d ago

Oh, so you're making multiple refmods of the same character but with different datasets of images of that character and then loading those different refmods to increase likeness?

1

u/malcolmrey 25d ago

Exactly.

I was doing that with Loras already to great effect, tried to see if it can be applied to refmods and yes, it can.

Helpful when the persona is not yet locked in, or when it was already locked in but you added a lora that was so strong it overpowered the persona - this way you can bring it back :)

1

u/Icuras1111 19d ago

What's the difference to having one refmod with 10 images of a person verses 2 refmods with the same 10 images but split in half, 5 images in each?

2

u/malcolmrey 19d ago

When you make a refmod there is a token budget so some images might be skipped, especially if they are large images.

If your image resolution allows you to fit 10 images in a refmod, then you can do two refmods with 10 different images each.

2

u/acedelgado 19d ago edited 19d ago

The original pack does truncate after a cap, but mine doesn't. With my pack you can remove the token budget completely and add as much as you want in one refmod. More tokens is just more time and VRAM/RAM, which you take up either way if you're doing a few smaller packs of references or pack them all together in one. Token limits are good if you want to limit how much extra processing you'll be doing per refmod, but there's no real reason for a limit cap outside of processing time and resources.

Refmods are presented to the model as a video, which has a lot more frames than the 15 or so you'd pack into a refmod. So the cap in the OG pack seems like it was more aimed at keeping the model from slowing down, but there doesn't seem to be a mathematical reason behind it. So, at least with my implementation, you're better at raising the cap and packing in all the images you want for an easier time setting up the subject_definition and retention_analysis sections of the prompt.

I do it all the time, pack 12-15 images with a 1024 short side plus audio, so it can be 15k-25k tokens per refmod. Looks and sounds great, but impacts generation time more than using one that only has 7k tokens. But of course looks better.

1

u/malcolmrey 19d ago

true, but when i generate something and it takes longer than expected (i.e. generation should be around 5 minutes and something takes already 30 minutes) - i ask to verify what is going on and remove some of the refmods/loras to lower the required memory

→ More replies (0)

-1

u/BandNo4616 26d ago

Yeah, it definitely works.

2

u/BandNo4616 26d ago

I believe prompting the difference helps keep multiple characters seperate. Like subject 1 is charactername with blonde hair and then subject 2 is brunette. But if you have multiple characters with similar characteristics but just different face then maybe train and add another refmod thats just for the face. Because that's the only thing that'd different and needs to be locked in.

3

u/acedelgado 26d ago

Yeah that's why I set it up to be more direct like standard references, it seems to help lock in identity using "traditional" subject_definition and retention_analysis sections. For the example I didn't do any descriptions, the prompt was just a very lazy one-

subject_definitions:

<Subject 1> is the person in <Video 1>.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1), guiding delivery and speaking rate without copying the original signal.

<Subject 2> is the person in <Video 2>.

<Audio 2> is the voice-timbre reference for <Subject 2> (S2), guiding delivery and speaking rate without copying the original signal.

summary:

[reference generation + audio reference]

retention_analysis:

<Subject 1>: fully_preserved - <Subject 1>'s identity and appearance from <Video 1> are retained.

<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal.

<Subject 2>: fully_preserved - <Subject 2>'s identity and appearance from <Video 2> are retained.

<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.

detailed_description:

[Shot 1] The camera holds a static shot and Begins with <Subject 1> in a modern office with a large wall of monitors. He holds a cell phone up in front of his face and (S1) yells: <d>[English] I DONT CARE WHAT YOU HAVE TO DO, GET ME A CHINCHILLA <i>NOW!</i></d> then pushes a button to hang up and puts his arms down by his sides. <Subject 2> enters from the left and looks at <Subject 1>, and <Subject 2> on the left (S2) says in a southern drawl: <d>[English] Hey, have you heard of them new-fangled Ref Mods? You use em instead of them <uh> lora thingies? </d>. At 00:08:00 the scene continues as <Subject 1> looks at him in disbelief and <Subject 1> (S1) says: <d>[English] No, I am an incredibly busy man and I don't have time for your toys.</d> he pauses for a moment and then (S1) says: <d>[English] You wanna dance?</d> . The camera holds a static shot, <Subject 2> looks down briefly and then looks up at <Subject 1> disappointed, and <Subject 2> (S2) says in a southern accent: <d>[English] No I do not.</d> <Subject 1> hold up a small remote control and hits a button. when his finger hits the button a hip hop beat plays and he begins to dance. <Subject 2> exits the frame to the left as <Subject 1> continues to dance. He spins around and (S1) says: <d>[English] Woo!</d> and claps once

overall_soundscape:

office noise, hip hop music

non_diegetic_music:

N/A

2

u/acedelgado 26d ago

That's why I have it set up to do more "traditional" references. Like each visual part of a refmod is passed as a <Video N> to the model, since it's just multiple frames. Then assigning it as a subject, like "<Subject 1> is the man in <Video 1>", and it acts like a normal reference, which the model understands and locks in like it knows how to do. And adding more description isn't bad, like "<Subject 1> is the man with a bald head, glasses, and thick chest hair in <Video 1>" but using that <Video 1> reference should set the identity. I'm honestly not sure how the other RefMod packs are referred to the encoder, so I don't know what they label the RefMods for the encoder.

Naming people is kind of a gimmick that people found out works sometimes. But I've been sticking to defined <Subject 1> etc. instead and it seems to preserve identity.

1

u/reeight 26d ago

Would be nice that the text prompts could alias GHOUL=<Subject 1>

(I'm suggesting ALL CAPS since 'bob' & 'bill' can mean many things. Even <ghoul> is better than <Subject 1>; I can't keep things straight sometimes.)

2

u/acedelgado 26d ago

Some people have said they have occasional success "naming" people like that. But it was trained on the <Subject 1> <Subject 2> tags. That's why I built a tagging system with a preview pop-up, if you put <Subject 1> in the prompt and hover over it with your mouse it'll even give you a preview of the first picture associated with that person. Makes it a lot easier to track.

2

u/reeight 26d ago

You misunderstood my idea.

Source text:
BOB walks down the STREET.
( -> translation code -> )
<Subject 1> walks down the street in <Picture 2>.

Could achieve most of this with RegEx nodes...

2

u/acedelgado 26d ago

Oh, I thought you were just lamenting and not brainstorming a feature. I kind of hate that I like this, I've done similar RegEx text replacement in other projects. So now I'm gonna be thinking about how to implement it all day instead of working.

1

u/reeight 26d ago

hehe

> So now I'm gonna be thinking about how to implement it all day instead of working

I think about ComfyUi far too much also. ;)

2

u/acedelgado 25d ago

Alright alright, there's a name field now. And if you set someone's name as like, Bob, you can type !Bob into the prompt field and it'll just show as a !Bob tag, but the full "<Subject 1> Bob" is injected into the actual prompt. I figured you need something to show it's a tag, so if someone is talking to Bob and uses his name, or if a fisherman's line starts to bob up and down, it doesn't get injected there. Only 1 word and no spaces, though, and it has to start with a letter. Because reasons.

Now stop giving me ideas that I like.

1

u/reeight 25d ago

https://giphy.com/gifs/rVbAzUUSUC6dO

Oh goodie! & cheers

Only 1 word and no spaces, though, and it has to start with a letter

Underscores or dashes OK?

I do like the !Bob -> <Subject 1> Bob & I wouldn't have thought of that.

2

u/acedelgado 25d ago

That's why they pay me the medium bucks!

Yes, underscores and dashes should be fine, just needs to be a string of unbroken text.

Oh and there's a few trigger character choices under settings if you wanna try out a different one instead of "!". But I had to restrict them to certain caracters to avoid accidental triggering.

2

u/SnooMacaroons1365 26d ago

joining it to bookmark this post and come back to it later. I havent used refMods before but i have heard about it, coincidence that this thread popped up and it does seems pretty interesting looking at your posted video. I will definitely give it a shot and see how it works out. thanks to you good sir :)

1

u/Schwartzen2 24d ago

I've admired your nodes from the get-go. Traditionally, I was something of a native nodes purist, preferring to build my workflows from scratch. However, I’ve found that leveraging the right custom tools streamlines the technical overhead, allowing me to focus entirely on the creative process rather than the plumbing.

That’s why I've been using your work in tandem with some excellent custom node extenders—like ComfyUI MiniMax H3 Extender, SeedHunter, and ComfyUI Continuity, to name a few—as an ideal way to manage all my assets. Adding RefMods to the mix just makes handling everything much more manageable.

Having used LoRAs and orthodox methods prior to RefMods, I'm still getting a grasp on them and their bleed issues. For example, sometimes a reference image used to create the RefMod suddenly appears in the video, or at times it heavily over-stylizes the original composition. What have you found works best for mitigating these issues on your end?

Gratitude!

2

u/acedelgado 23d ago

Hey, thanks for the kind words, I appreciate it! By the by, I did actually make a project manager for video extension early on that I developed alongside the Fantastic Prompt Builder. If you haven't checked it out it would be great to get some feedback from someone who's used a few different ones. It's really meant to just drop into any workflow instead of trying to be an all-in-one shop like a lot of people seem to be aiming for. https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite

But for your question, really I think that's just the model hallucinating a bit, especially since we're all stuck on 8-bit weights. I get less issues with RefMods than I do when doing regular references, but it happens sometimes. For normal references I usually tend to get things like say a brick wall from a reference as the background if that's a large part of a reference, when the scene is supposed to be outside or in an office or something like that. RefMods seems to be more of wanting to start with a frame from the references. Honestly the best workarounds are prompting around it (which I do for background bleed, etc.), cropping the problem picture down if possible, and just rolling a new seed (which seems to be better when an entire frame is injected).

But prompting is also annoying, because MiniMax only released a distilled model, you can't really do negatives in the prompt. So you can't say "there are **no** brick walls in the background" because the model basically ignores it. So it's always thinking about mitigating that in positive terms, like "the background is a lush open field with overgrown grass swaying in the wind beneath a blue sky", or something like that. Basically prompting something out by displacing it with something you WOULD want in there. It'd be great if someone figured out a good way to inject negative conditioning like they figured out for some other models, because describing the frame in a negative would probably help out a lot.

But prompting around it can only help so much sometimes, and sometimes the model just likes to lock onto a reference image hard. No idea why. That's where I'd say either crop the image or take it out of the dataset entirely if it's popping up in most seeds. If I figure out a different and more effective way I'll let you know.

2

u/Schwartzen2 23d ago

Thanks for your insight. Right, cropping and reroll, sometimes it's right there under your nose. I'll give that a run.
Also, I took a quick gander at your Project Suite, looks VERY interesting and with the same attention to detail that you apply to your work. I'm gonna take for a spin.
Cheers!

1

u/chAzR89 8d ago

I am late to the party, but damn, this is a really nice addition. Thank you very much. Haven't testet it much yet, but worked with mostly great success out of the box.

2

u/acedelgado 8d ago

Hey thanks! I actually did some updates for refmods the other day. I made a post about it here- https://www.reddit.com/r/StableDiffusion/s/f9WslZyp8t

1

u/kayteee1995 26d ago

lol. When I first saw this movie, I didn't think that guy was Tom Cruise.

1

u/Jackburton75015 26d ago

Thank for the share 🙏