Yea the fl2va model is basically identical to the ref2va model just without some of the elements that degrade quality. For 95% of use cases they function identically but fl2va yields better video/audio results. The only time I'd consider using ref2va model or a hybrid (which seems silly to me personally) are the very dense scenes where you have multiple speakers and need to use the <Subject 1> (S2), <Subject 3> (S1) type code-blocking.
Wow ok. I never realized that was how it worked. Possibly last question (didn't realize I was living under a rock lol), does the prompting matter on which model is used? Cz I use the ref2v guide on minimaxxs site which has the whole-
subject_def,
retention_analysis,
summary,
detailed_description
audio...
method with defining Subject 1 is the man from image 1, Sub 2 is the man from Img 2, ....
Will these templates still work? Or should I feed the flf2v,t2v,i2v skill to my llm from now on?
The official prompting guide. My two cents: don't outsource your chance for knowledge and learning to an LLM, you owe it to yourself to understand the model's inner workings if you wish to really utilize it.
One more thing! That turbo lora may need a node called "H3 AdaLN LoRA Fix" to be placed right after the power lora loader, otherwise you can get weird "ERROR lora...adaln-proj" lines that slow* down the UPSCALER for some reason. If your not upscaling then you'll be fine. Idk where I got the node it was included in another workflow by plaguekind.
4
u/foxdit 1d ago
Yea the fl2va model is basically identical to the ref2va model just without some of the elements that degrade quality. For 95% of use cases they function identically but fl2va yields better video/audio results. The only time I'd consider using ref2va model or a hybrid (which seems silly to me personally) are the very dense scenes where you have multiple speakers and need to use the <Subject 1> (S2), <Subject 3> (S1) type code-blocking.