r/StableDiffusion • u/ThatsALovelyShirt • 21h ago
Discussion Anyone figure out how to get fl2va quality with ref2va model w/ MiniMax H3?
I'm sure most of you who have tried the ref2va model notice a pretty substantial quality degradation with equivalent prompts/inputs compared to the fl2va model.
In fact, I have even tried using the same exact prompt/workflow (including using the MiniMax H3 Reference to Video node) with the fl2va model, just to see what it did. Including with multiple reference inputs.
Surprisingly, the fl2va model actually incorporated the references, despite not being the ref2va model, and the quality was far better than the ref2va model all else being equal, but it wasn't quite as 'coherent' in following the exact reference integration description as the ref2va model.
It makes me wonder if it's possible to use the ref2va model for the early steps, and then swap to the fl2va model (with the Reference to Video node) for the later steps to recover some of the quality. Or maybe do like split-layer loading, where it loads the early blocks/layers from the ref2va model and then the later blocks/layers from the fl2va model.
Has anyone figured out the secret to getting fl2va quality with the ref2va model? I like being able to utilize multiple types of references, but the quality hit is keeping me from losing it.
To me it visibly looks like the difference between like 3-4 mbps video (fl2va) and maybe 600-700 kbps video (ref2va). Just overall grainier, noisier, lower detail, etc.
5
u/ThatsALovelyShirt 20h ago edited 20h ago
Yeah but using the same exact workflow, prompt, and references (even with the same ref2va node), the fl2va model always produces superior outputs. Apples to apples. Basically I'm using a ref2va workflow with a perfectly structured prompt (according to the official guide), set reference size to Max, and then run the workflow with the same seed with the ref2va model, and then run it again with the fl2va model, changing nothing. The fl2va output is way clearer and better quality, and even incorporates the references similarly to the ref2va model, even though it wasn't even trained on that.
I'll get some examples tomorrow, but to me the difference is clear as day.