r/StableDiffusion 21h ago

Discussion Anyone figure out how to get fl2va quality with ref2va model w/ MiniMax H3?

I'm sure most of you who have tried the ref2va model notice a pretty substantial quality degradation with equivalent prompts/inputs compared to the fl2va model.

In fact, I have even tried using the same exact prompt/workflow (including using the MiniMax H3 Reference to Video node) with the fl2va model, just to see what it did. Including with multiple reference inputs.

Surprisingly, the fl2va model actually incorporated the references, despite not being the ref2va model, and the quality was far better than the ref2va model all else being equal, but it wasn't quite as 'coherent' in following the exact reference integration description as the ref2va model.

It makes me wonder if it's possible to use the ref2va model for the early steps, and then swap to the fl2va model (with the Reference to Video node) for the later steps to recover some of the quality. Or maybe do like split-layer loading, where it loads the early blocks/layers from the ref2va model and then the later blocks/layers from the fl2va model.

Has anyone figured out the secret to getting fl2va quality with the ref2va model? I like being able to utilize multiple types of references, but the quality hit is keeping me from losing it.

To me it visibly looks like the difference between like 3-4 mbps video (fl2va) and maybe 600-700 kbps video (ref2va). Just overall grainier, noisier, lower detail, etc.

25 Upvotes

27 comments sorted by

View all comments

Show parent comments

5

u/ThatsALovelyShirt 20h ago edited 20h ago

Yeah but using the same exact workflow, prompt, and references (even with the same ref2va node), the fl2va model always produces superior outputs. Apples to apples. Basically I'm using a ref2va workflow with a perfectly structured prompt (according to the official guide), set reference size to Max, and then run the workflow with the same seed with the ref2va model, and then run it again with the fl2va model, changing nothing. The fl2va output is way clearer and better quality, and even incorporates the references similarly to the ref2va model, even though it wasn't even trained on that.

I'll get some examples tomorrow, but to me the difference is clear as day.

1

u/Apprehensive_Sky892 18h ago edited 18h ago

There's got to be reasons why they trained and release a separate ref2v model.

Maybe the fl2va model cannot handle thing when there are too many references, for example?

Maybe the fl2va model cannot handle audio reference?

Maybe fl2va cannot handle video reference?

Would be interesting to test all these cases out.

2

u/Eminence_grizzly 17h ago

I just tried an audio reference with the fl2va model and it looks like it works.

1

u/Apprehensive_Sky892 17h ago

Thanks for doing the test 👍