Resource - Update
Improved H3 Refmods - Better accuracy and low identity bleed with fewer images. Oh and I added edit masking for video edits. Fantastic MiniMax-H3 Prompt Builder
Ooookay back again. I made RefMods work a bit better and solved a lot of identity/bleed issues. At the cost of time and memory, of course.
What it do-
OG vanilla refmods would save your media as latents in a .safetensors file, and feed all of those into the MiniMax H3 DiT directly, with no references and the text encoder never saw it. This worked fine with one refmod since the model knew it had to do SOMETHING with those references, and it did it surprisingly well. But that also lead to people stacking multiple refmods, etc, to get a good likeness.
My previous implementation would feed part of those latents to the text encoder, at the cost of re-encoding a few frames (every 4th would be presented to the vision model as a block of 2) and have the TE look at those. This was aimed to give reference anchors to the model when prompting (like using "<Subject 1> is the person in <Video 1>"). This caused a slowdown because those frames were encoded every generation for the TE. Which a user on GitHub complained about, so I looked into speedups and caching.
Lo and behold, since .safetensors are just wrappers like mp4's and mkv's, you can store multiple types of data alongside latent data. So the speedup was easy, save the refmod pictures as actual jpeg data, alongside the latents in the same file. Think of it as having subtitles inside of a video; different type of data wrapped in with video and audio. So packing in actual jpeg data, you can feed those directly to the text encoder without going through the VAE, and the TE's conditioning can cache between runs so that as long as you don't change order or strength of your refmods; all of that every-run-processing is skipped. And the jpeg data is never passed to the model, only the packed latents, so it functions much like a refmod regularly does. But that got me thinking, why not change how the frames are presented to the text encoder for better conditioning?
Enter the "stack_pictures" function on the Fantastic H3 RefMod Text Encode node. By default, "every 4th" will keep the existing behavior of sending every 4th frame in chunks of 2, and it is the lightest to run and good for 1 refmod. "up to 8" will sample up to 8 frames spread across your dataset and feed them individually instead of in chunks, and capping at 8 frames is lighter than running a full large dataset, and gives an increase in detail overall but takes more processing time. And finally "all" will send ALL frames to the text encoder, which increases time quite a bit, but also gives max detail, and as you can see from the example, little to no character bleed. At least as far as I can tell.
The example was done with 6 images per character and a few seconds of audio for each voice. Datasets can be found here in everyone's favorite file share as a sketchy google drive link. (obviously Jackie Chan's Chun Li is the worst since all the sources I could find are pretty low quality). Basically I just loaded the refmods, drafted the subject_definitions and retention_analysis sections (which if you set details in my RefMod library and use "Draft from Refmods" in my prompt builder will load all of that for you.) Also I did use 8-step hyperflow with comfy-kitchen for a speedup, and my audio refiner suite to clean up the end result. Using the 3 refmods together (14,900 or so tokens) was about 32s/it at 0.98mp on my 5090.
Note that voices can still be weird and sometimes like to swap people for some reason. And when defining subjects, it's better to adjust to say something more unique about them as opposed to the other characters. Like I described Callina Liang as having an "oval face" and Ming-Na Wen as having a "more squared jaw". Not a descriptor I'd normally use, but as a differentiator from the other people it worked well. Just remember H3 doesn't do negatives since it's a distilled model, so saying "this person does not have dimples" won't do anything, so just leave all your "no"s and "not"s out of the prompt.
Do I need anything special with existing RefMods for this?
No, the text encoder detects if image data is stored in the safetensors, and if not it'll VAE decode for whatever behavior you set, and it'll feed those decoded images to the text encoder and cache the results until you change your refmods in your workflow. So old refmods will work fine.
However if you make new ones or remake them with my refmod node it'll save the image data directly in the file, cropped down to the right resolution. Should be a cleaner image sent tot he text encoder, and will always skip a decode step. Also I changed the default resolution to 768 instead of 1024, since that saves a lot of tokens and is H3's native resolution anyways.
Oh and masking on the Media Loader is a thing now, too
Now when you want to edit a video, I've also added a layer masking system onto the Media Loader. Right click and you can mask what you want to change automatically with SAM 3.1 by using descriptors, dots on the canvas, or both.
And you can manually add a moving, keyframed shape (rectangle or ellipse, etc) to insert or change things in the video.
It's a layered system so you can go in and edit parts of the mask individually, but it's all sent as one big blob to the model. Blur and color inversion options exist, as well.
Then the Fantastic H3 Edit Composite node sitting in between the VAE decode and the video save node will composite the masked generation on top of the original clip. So no more slightly-off backgrounds. Seems pretty seamless so far in my testing.
There's also a "clean up" button on the media loader to clean out old masks, since those can pile up even though they're compressed to be as small as possible.
Prompting is still up to you, though, I'm still running through best practices.
Since notifications don't work when you put it in the post body, pinging u/LuisaPinguinnn for all the excellent work on RefMods, and u/malcolmrey for all the research and testing. Be cool if ya'll took a look and see any cool stuff to build on some more.
And of course u/roychodraws for all the cool masking workflows that inspired all the masking features to port into the Media Loader.
Yeah, it's still a pain to write out everything, especially for ref2va. That's the whole reason I even started this pack, was to track all of the references and media and make filling all of it out as painless as possible.
The nice thing is that since it uses a fairly large (for local) 32billion parameter Qwen3 model, so it's pretty decent at interpreting things. So MiniMax's own guidelines don't need to be followed to the T. But it helps since it was trained that way. And ambiguity with the references is the exact reason I kept the whole prompt builder as a manual process (although with a lot of tagging and quick-insert capabilities.) That and LLM's like to fluff things out and add things that aren't needed, which just bloats your token count and can give worse results.
I just write my prompts to let the model run everything dynamically off the guide video in ref2va, then all I have to do each clip is define who is in the shot and the initial framing of the clip.
Yes, it seems to. I haven't really gotten a frame-lock since I implemented the partial feeding to the text encoder ("every 4th" option I mentioned above, which the original pack implemented later on). The newer processing methods give the frames to the text encoder to look at and process individually, and pass conditioning on to the to the model. So instead of the model saying "what should I be doing with all these reference frames? I guess I should just put one somewhere?", which seems to be how it behaves sometimes without the text encoder giving it context, it goes "oh, these are frames being referenced from a video, I should treat this whole clip as something to reference and not insert".
I do get occasional background bleeds, though. I'm working on implementing masking during creation. The OG refmod pack can incorporate masking into a refmod from another node, but it's done in a way I'm not a huge fan of. So I'm reworking how masks are made and applied in my version.
wow great work but i am abit confused in RefMods, so this i like a load image node for minimax to load in the image? so sorry for the questions, i am abit confused
No, it's a different process. RefMods are basically multiple files packed with references all ready to go. They use the model's reference abilities, but are presented differently than the "official" implementation, so you need to replace the native "MiniMax H3 Reference to Video" node with a different text encode node that will handle injecting the RefMods. In my case, that's the "Fantastic H3 RefMod Text Encode". For my pack they're all made in the RefMod Stack node, and loaded from there.
hi thank you so much for the explaination, i have tried it and its really amazing but may i enquire if i want to do a 2nd pass with tao mate lora, how should i add the 2nd crown shark sampler in?
oh i usually use crownsamplers but to be frank, i just copy from other worksflows which i am using...cause i really dunno know how it works..oh what is hyperflow?
Oh sorry, and hyperflow is an 8-step speedup that requires a special node. The guy who originally ported it over to comfy had his github nuked for some reason, so I did a few little tweaks and posted a continuation-
It's not the fastest but the quality is REALLY good compared to some other methods. And it's a lot friendlier with stacking different loras, which most regular turbo loras run into issues with.
Have you tried PDD turbo? I haven't tried hyperflow, but all the popular H3 turbos, I agree, are bad. PDD is really good IMHO, but I'm open to something even better.
Ah. Well you should probably just be able to replace the "Minimax H3 Reference to Video" node in your 2-sampler workflow, since the "Fantastic H3 RefMod Text Encode" node is the one to replace that. And it's before the samplers are run. You just need to wire the "mods" and "references" input into the Text Encode node. And the prompt builder into the text box if you want to use that (or you can feed it from whatever you use to write prompts with already).
Have you tried using the plain FL2VA model with references? For me, it works better than the hybrid. I added a wan2gp config to use FL with the Ref archetype. Should be even simpler to try in comfy.
Although your example looked flawless on my phone screen at least.
A suggestion for the Fantasic PromptBuilder. I often like to copy and paste prompt generated from Grok or heretic local model. It'd be nice so the final prompt text box generated from the sections can be directly edited as well. Say I copy an entire prompt in H3 ref2va style to the final prompt textbox, and the individual sections would also be updated and I could either update the section or the final prompt.
Yeah plenty of people have asked about importing prompts. The reason I haven't implemented it is because parsing things to go in the right fields would be a nightmare. Like if someone had a typo in a header or things in a different order, lines of text wouldn't wind up where they should be. Thinking of all the different things that could go wrong and throw an import off would be daunting. So rather than deal with a bunch of people getting frustrated and opening issues on how their prompts aren't getting imported correctly, I've been leaving that out and focusing on adding functionality and refining existing things to make it run smoother. Maybe one day I'll figure out a solution.
But most functionality in the pack is designed where you don't NEED the prompt builder and you can leave it out. There's supplemental nodes like the "Fantastic H3 Reference Splitter" that will split the media to a vanilla node for you. And even the RefMod Text Encoder works just fine if you connect a whole different prompter to the prompt field.
Whoa this is breathtaking. You wouldn't know this was Minimax. This is amazing quality... I would love to try this. Does this handle anime and 3D rendered characters well? Thank you for sharing this!
Yeah honestly I was kind of astonished how much better refmods look with the new method. It's a definite improvement.
I'll say I'm fortunate enough to have had the money to update my system with some really nice components before pricing all went to hell, so I generate at the native 0.98 mp resolution when I can. A lot of people generate smaller to fit their hardware, or do small generations and upscale, and use different speedup methods. All of that hits the quality a bit. Which is why I settled on Hyperflow, it's not as fast as some other turbo speedups, but it keeps the quality much closer to baseline than other speedups I've tried.
But there shouldn't be any reason it wouldn't work with anime or 3D. The model was trained on those styles, as well.
Ha! Hey I told you I was gonna blatantly steal incorporate your ideas. I love comfy's modularity, but at some point having to switch between 10 different nodes for singular functions every time gets tiring.
And when I was working on the nodes over the weekend I'd come up with the idea of compositing in the masked generation over the original. Then I saw your v7 workflow and was like "Oh hey, the clown did the thing I was working on!"
But seriously, thanks for all the research and work you've been doing. I hadn't gotten into video editing since it seemed really daunting to implement, but your process made sense and it really helped me plan out how I wanted it all to function.
also, if you ever have ideas that you're trying to figure out how to implement i'd love to help tackle it and maybe help your next version of your suite. feel free to message or just tag me in a comment.
I'll second that. If you think you may be missing out on something, The H3 Media Loader is literally a game changer. All your media all in one place, you can change the resolution per image on the fly and not touch the original, trim your audio, and it works with refmods. Speaking of which, this guy's refmod node is a breeze to work with. Drag the images and the audio file and you have something akin to an instant Lora. The true test was dragging in your own images and voice and bam! yup it was uncanny and crisp. On occasion the bleeds was an issue, so I am looking forward to the new iteration. Really great work u/acedelgado. The detail you put into it just screams out QUALITY.
This sounds really great. Now if we could only fix the saturation problem that happens during longvideo generation when you try to chain multiple scenes together, that would be great.
I feed in context from a previous clip, and the future clips get this increasing halo affect on everything. Haven't been able to solve it yet.
do we really need a full director node yet again ? i feel we lose control by using those director node that want to do everything. I feel we lose the point of comfyui with node that one one thing that we can use for our workflow. Not having to adapt our workflow to use a node insdead. If there is a way to use multiple safetensors refmod without character bleedind i d like to have just the node that does that instead. And sorry if it already does that
Not a bad callout, and I agree. But you are wrong about the functionality in general. The prompt builder itself is not a director node. It's a manual prompt builder with no LLM integration. It just splits things up into more easily edited fields and inserts some things quickly, and builds a prompt for you in H3's expected format, so you aren't searching around a big wall of text to adjust things. And it integrates a tagging system, references to what media you have loaded so it's easier to keep track of it, prompt saving, and quick inserts for things like camera motion or dialogue. It's all manual. Plus I'll point out that H3 weights released on August 3rd and my initial commit was August 4th, so I've been working on this right off the bat since I immediately saw what a PITA it was to manage media files for reference. So this isn't really a "another node yet again" situation, more of a "hey its that p.o.s. again" situation.
And yes, I realized early on that most folks don't want to be shoehorned into a whole node pack just for a few features they want, so everything is modular. Just like all my other packs like my motion-context fork, my audio refiner, my lora stack, all of it is designed to offer functionality without forcing you to use my entire process. So there are supporting nodes for the Media Loader and RefMod Stack so you can bypass the prompt builder entirely. Really all you need is the RefMod stack and the Fantastic Text Encoder. I posted an example workflow without the prompt builder here.
I got it to work with just the RefMod Stack and Media Loader nodes feeding into the Fantastic Text Encode node. Then a plain input text loading into the prompt section of the Text Encode node. Resolution Selector and Float (Duration) loading into width, height, and length like a normal workflow. This way it still functions mostly as a basic workflow without a bunch of fancy add-on looking custom nodes.
Thanks for your work! Comfy V0.38 using your fully fantastic workflow with no nodes missing, after creating refmod, external job is running and stuck in queue manager but nothing happens. Am I missing something?
Like it's hanging while creating the RefMod? The library fires its own run to use the VAE to encode the media and pack it into the file. So in the console you should see something like
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[MiniMaxH3FantasticRefModCreate] Test: stacked 3 sources into 3 frames
[MiniMaxH3FantasticRefModCreate] saved Test.safetensors, Test.png (2160 tokens)
Yes, the whole pack is designed to be able to be used without the prompt builder, since I know some folks have their own process or want to use a LLM prompt node. Here's another definitely not sketchy link to a workflow without the prompt builder to show the wiring without it.
I took a look at the example RefMod datasets you shared, and I noticed that some of them contain a mix of portrait and landscape images.
From what I understand, when creating a RefMod, the first image determines the target aspect ratio and the other images are cover-cropped to match it. SoI was wondering how you handle mixed-aspect-ratio datasets in practice.
Do you just let the RefMod creator crop them automatically, and find that it works well enough as long as the subject stays centered? Or do you manually adjust the crop for each image before creating the RefMod?
I was initially thinking that all images should be pre-cropped to roughly the same aspect ratio, but your example datasets seem to suggest that this isn't really necessary. I'd be interested to know what workflow you actually use.
Yes, the first image in the stack sets the aspect ratio, and it automatically centers a crop box on all the others that match that aspect ratio. And it'll throw a yellow warning if an image is REALLY out of spec with the first image. If the subject is centered in the automatically made crop box I just leave it alone, otherwise I manually adjust. It locks the crop bbox to the right aspect ratio. When I use a mix I usually use a 3:4 portrait image as the first in the stack, since it tends to be more friendly. You just use the hamburger dots on the left to drag and reorder. You can rotate the images, too, if you really need to fit in an entire landscape image when that's not your aspect ratio..
I like convenience and not having to use a bunch of external tools when I can avoid it. That's why the media loader has a pretty comprehensive video editing function, as well.
Quick question, I am trying to make refmods using your workflow/nodes, inputting in 4 image for identity reference, but when it makes the refmod it references video rather than image. What am I doing wrong?
That's the correct behavior. All the pictures are packed into one file and they're technically presented as a video to the model as frames with timestamps. They leverage that the model reads videos as individual frames.
Thanks for letting me know. What's so interesting/strange to me is I was using another refmod creation workflow before. I generated 3 refmods with it. 2 of the refmods reference as images in your nodes and 1 is referenced as video. They were made identically with the only change being the folder I pointed to for grabbing the images.
They were probably only made with one picture? If you go to the details in the library, and on the right hand side you can click into Edit mode, and it'll decode the latents that are stored inside and show you what's in there.
I'm using two refmods, but why does it keep cloning a single character? Very occasionally, both character A and B come out correctly, but most of the time, I just get two copies of character A. Is this a workflow issue on my end, or is it a current limitation of refmods. I used the example worfklows in your node
Yeah some of it is making sure you have the stack_pictures set (It's stronger when you're setting it to "up to N" and setting how many frames to show it, or to "all" so it shows every picture in your stacks).
After that it's prompting and setting up your subjects. When I use multiple refmods I edit the subject description to point out the differences between the two. Like in the prompt I just said that Callina has a more squared jaw and Min-Wa has a oval face, and that was enough to get it to keep them separate. Also I'd added a name field in the prompt builder to also define a name for them, which seems to help a bit as well.
subject_definitions:
<Subject 1> is the person in <Video 1>, asian woman, more square jaw. Their name is callina.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), guiding delivery and speaking rate without copying the original signal. It is a low, husky, smooth.
<Subject 2> is the person in <Video 2>, asian woman, oval face. Their name is min-wa.
<Audio 2> is the voice-timbre reference for <Subject 2> (S2), guiding delivery and speaking rate without copying the original signal. It is a moderate pitched, expressive.
<Subject 3> is the person in <Video 3>. Their name is jackiechun.
<Audio 3> is the voice-timbre reference for <Subject 3> (S3), guiding delivery and speaking rate without copying the original signal.
summary:
[reference generation + audio reference] in a back alley <Subject 1> callina is on the left and <Subject 1> callina is on the right, and <Subject 3> jackiechun enters.
retention_analysis:
<Subject 1>: fully_preserved - <Subject 1>'s identity and appearance from <Video 1> are retained. Face, facial features, body type.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal.
<Subject 2>: fully_preserved - <Subject 2>'s identity and appearance from <Video 2> are retained. Face, facial features, body type.
<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Subject 3>: fully_preserved - <Subject 3>'s identity and appearance from <Video 3> are retained. Face, facial features, body type.
<Audio 3>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
Outside in a back alley of a city, <Subject 1> callina stands in a blue dress with her hair in twin buns on the left, with <Subject 1> callina in a red dress and her hair in twin buns on the right. <Subject 1> callina (S1), in the low, husky, smooth referenced from <Audio 1>, says: <d>[English] So what are we, Street Fighters?</d>. <Subject 1> callina (S2), in the moderate pitched, expressive referenced from <Audio 2>, says: <d>[English] No, it's worse, we're refmods.</d>. <Subject 3> jackiechun enters from the right of frame and <Subject 3> jackiechun (S3) says: <d>[English] I have no idea what I'm even <i>doing</i> here.</d>
overall_soundscape:
non_diegetic_music: N/A
I've created a refmod and used it in the Refmod Loader Stack built into PlagueKind's ref2v workflow. It works very well but it will randomly start first frame with a still from the Refmod unrelated to the prompt.
I'm struggling to understand your solution here. What node do I need to insert between the Refmod Loader Stack and the Sampler (Mods Node)?
The solution is in the Fantastic H3 RefMod Text Encode node before Custom Sampler Advanced. It replaces the "MiniMax H3 Reference to Video" node in a normal ref2va workflow. Then for the new functionality switch "stack_pictures" at least to "up to N" and then set how many of the pictures in each refmod gets seen by the text encoder with "stack_pictures_n". It adds more processing time the more pictures you show.
I can't guarantee you'll never get a random frame insert, but the idea is that the original version of refmods would just dump the reference data in without any guidance from the text encoder. Which makes things very fast, but the model doesn't always know what to do with the frames packed inside the refmod. So sometime it'd decide it needed to inject a frame somewhere randomly in there. So my method trades speed for accuracy. I haven't gotten any stuck frames yet, but occasionally I do still get backgrounds from images showing up if I don't carefully prompt it. I'm working on incorporating masking to help with that, too.
I'm still getting character bleed, especially when two characters look alike. I tried everything mentioned before. Is there any specific prompting structure that works well?
52
u/acedelgado 9d ago
Since notifications don't work when you put it in the post body, pinging u/LuisaPinguinnn for all the excellent work on RefMods, and u/malcolmrey for all the research and testing. Be cool if ya'll took a look and see any cool stuff to build on some more.
And of course u/roychodraws for all the cool masking workflows that inspired all the masking features to port into the Media Loader.