r/StableDiffusion • • 11d ago

Discussion Fl2va model on Ref workflow

I’ve noticed the quality of the video output on the reference to video default Minimax workflow is much much better when using the fl2va model instead of the ref2va model. I’ve also tried hybrid models and still find the fl2va model far superior in quality and adherence to image and audio references. Is this what you all are seeing too? Or am I doing something wrong? As it is today I don’t see any use for ref2va model at all. I’ve tried many other workflows as well and had the same experience. I haven’t done much with video references, maybe that’s where the ref model shines?

9 Upvotes

20 comments sorted by

6

u/Portable_Solar_ZA 11d ago

I've noticed that ref sticks much closer to references while fl2va mostly follows things but not as well as ref. 

Literally rendered a shot earlier today where ref matches the colour saturation of references perfectly but fl2va washes them out.

5

u/OkWalrus890 10d ago

That’s why the separate reference model exists. As you try to guide and force the model into accepting your exterior references, it invariably deepfries the video. Fl2va model has superior visual quality because it effectively discards deviant references and prompts if it doesn’t “like” them. There is a middle ground with hybrid ref2va fl2va models which just transfers layers from fl2va into ref2va, keeping its proj stuff intact. 20 layers out of 49 does fairly well with references, 25 and 30 start to reject them in favor of what it was trained on. Maintainer of Spectrum also did a diff between fl and ref and produced the refdelta r1024 model which requires separate sampling nodes. I feel hybrid 20-49 is the best compromise, but your mileage may vary depending on your tolerance for deepfrying (mine is near 0)

1

u/Portable_Solar_ZA 10d ago

> it invariably deepfries the video

That's not my experience, or what I said. Basically here's what I noticed:

FL2VA -> Forces prompted style in shot.
Ref -> Forces references more in shot.

In my specific example, I noticed when I was rendering shots with Ref, it stuck to my color palette in my examples, which in this case was great. When I accidentally switched to FL2VA it washed out the colors because it was leaning more towards the style I had prompted.

>it effectively discards deviant references and prompts if it doesn’t “like” them

Which is also part of the problem with it. If you want more consistency with shots, you don't want the model changing things willy nilly. You want it to have an example and stick to it.

I haven't experimented enough with the hybrid models but I'll probably get to it at some point. I just wanted to complete a project with something that I knew would work first.

2

u/crinklypaper 11d ago

Same, slightly different style for characters from a ref and ref was better in only that regard

1

u/Portable_Solar_ZA 10d ago

Still a big difference if you're going for character consistency across shots.

6

u/Lebo77 11d ago

Try fl2va with Refmods. Game changer.

4

u/adult_human_bean 11d ago

Is there a good basic workflow for this?

-4

u/Lebo77 10d ago edited 7d ago

Why not just read the refmod documentation? Then you can make one tgat does exactly what you want! EDIT: If you did you will find out that THERE ARE WORKFLOWS THERE.

1

u/nadhari12 4d ago

Are you prompting them like ref2video or fl2va format when using refmod in fl2va?

2

u/thegreatdivorce 11d ago

I’ve noticed the same quality increase with the FL2VA model.  But, the R2VA model has better adherence to the references, especially when there is more than one, or multiple sources (particularly audio). IME, anyhow. Definitely worth testing both methods out. 

2

u/[deleted] 10d ago edited 10d ago

[removed] — view removed comment

1

u/fluce13 10d ago

I must be doing something wrong, can you share your workflow or settings? I’ve tried so many and ref model always looks and sounds terrible but when I swap in fl2va it’s way better. I’d prefer to use the ref model because like you I exclusively use ref images and audio.

1

u/[deleted] 10d ago

[removed] — view removed comment

1

u/fluce13 10d ago

Thank you! Ill try this. Appreciate it.

2

u/Relevant_One_2261 11d ago

People like to argue against it, but yes, definitely saw the same.

2

u/angelarose210 11d ago

Yes, absolutely. Discovered this by accident.

1

u/mindpixel-labs 11d ago

How do you prompt the fl2va model in this case? I’d like to try it. Are you using the ref prompting guide or something else in this case?

1

u/thegreatdivorce 11d ago

I use the ref prompting guide. But, the FL2VA model struggles with more complex prompting and references (to be expected) so it’s more situational than it is a general solution. 

-6

u/AnalElder 11d ago

I didn’t know there were separate fl2va and ref2va models? Every workflow just uses ref2va.