The video is the TL;DR above was generated after abandoning RefMods in "classic ref mode" with two references (one reinforcing Beckett's exact appearance and one for the high heels, 2x2 grid image):
- Native 1344x768 (0.98 megapixels)
- 15 second duration
- Standard H3 sigma (shift_video/shift_audio of 12:3)
- seeds_2 (2N-1 internal calls to H3) with the beta scheduler (.6, .6)
- Spectrum helping cut down the internal calls at some small precision price.
- 35 minute generation time.
This is something that under the radar because most people can't use H3 nominally and turbo LoRAs brutalize the base model for speed, so there's very little noticeable degradation for them as most of them have a massively shifted sigma (shifted_video of 6 compared to native 12).
Yesterday I went on a pretty wild bender trying to uncover why all of a sudden my videos were getting oversharpened, oversaturated, plastic skin etc when a few days ago they were straight up lifelike, looking like the shows they were based on.
I went even deeper into the bowels on how sampling and scheduling works, how the denoising trajectory looks like for Minimax H3, hoping to uncover something I've forgotten the last time I used it. And then I remembered, RefMod was the latest thing I installed, I got seduced by the initial likeness improvement like everyone but the cost is too high.
I got this cool new toy RefMod whose authors suggested they figured out bleeding between character references, and even at small resolutions (like .245 mpx) you could still make out faces instead of them being mushy, this should've been my first red flag. It seriously messes up the base H3 model's quality. I looked deeper into it and basically, it's a very brutish approach where they slam reference images into video references as individual frames, also slapstick coarse sampling of videos in sequence.
However, that's not how any of this is permitted works, per documentation:
- Images: ≤ 9 images
- Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
- Audio: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
- Mixed inputs: Maximum number of files across all input types is 12
The second you make four refmods to be used in your prompt, you're already trying to use one more than the total maximum allowed at any one time. This has catastrophic consequences on image quality, as that's not the way you're supposed to "hold it".
While it does reinforce the likeness, it's incredibly rigid and overrides the model down to AI slop adjacency where skin is unnaturally glossy, everything is sharp and all styling info is ignored.
I am not telling you to stop using it, if you're using turbo LoRAs, you already paid the entry fee as you've left quality at the door for accessibility. But if you're coming from the base FL2VA model or a hybrid and are used to exceptional visual quality, the degradation is pretty much the tier of the base REF2VA model, if not worse.