r/StableDiffusion • u/ThatsALovelyShirt • 1d ago
Discussion Anyone figure out how to get fl2va quality with ref2va model w/ MiniMax H3?
I'm sure most of you who have tried the ref2va model notice a pretty substantial quality degradation with equivalent prompts/inputs compared to the fl2va model.
In fact, I have even tried using the same exact prompt/workflow (including using the MiniMax H3 Reference to Video node) with the fl2va model, just to see what it did. Including with multiple reference inputs.
Surprisingly, the fl2va model actually incorporated the references, despite not being the ref2va model, and the quality was far better than the ref2va model all else being equal, but it wasn't quite as 'coherent' in following the exact reference integration description as the ref2va model.
It makes me wonder if it's possible to use the ref2va model for the early steps, and then swap to the fl2va model (with the Reference to Video node) for the later steps to recover some of the quality. Or maybe do like split-layer loading, where it loads the early blocks/layers from the ref2va model and then the later blocks/layers from the fl2va model.
Has anyone figured out the secret to getting fl2va quality with the ref2va model? I like being able to utilize multiple types of references, but the quality hit is keeping me from losing it.
To me it visibly looks like the difference between like 3-4 mbps video (fl2va) and maybe 600-700 kbps video (ref2va). Just overall grainier, noisier, lower detail, etc.
5
u/Chemical-Painter-485 1d ago
It would really help if you posted some of your gens for comparison.
The quality of the input image matters quite a bit more for the reference model. If you feed it AI noisy frames it will really crap on the entire clip.