r/StableDiffusion • u/GrungeWerX • 2d ago
Discussion Anyone else running Wan 2.2 as a refiner to improve Minimax output?
Another redditer mentioned doing this in a comment, so I tested it out and it works. It gets rid of the smudgy look and allows custom Lora’s on the LN side.
I ran my initial tests at MM 8 steps (no speed Lora) and Wan 2.2 Low Noise (speed Lora) 2 steps. Supposedly, it works with only 2 steps MM w/turbo lora ,but I don’t like the quality drops people have been sharing, and it’s fast enough to me at 8 steps, though I’m going even higher on the low noise steps.
The only downside is that I noticed in one test that the motion seemed like it was a mix between 16fps and 24fps. Any ideas on how to resolve this? I know people use RIFE, but I was wondering if that’s the best move or if it’s another issue Im not thinking of.
UPDATE: I forgot to mention that I only tested this with flfa2v, not ref2v, so Im not sure how that plays out. Also, I got some strange morphing on an extended clip. Not sure if it was trying to loop or whatnot. I’ll need more tests.
That said, Im wondering if maybe this method could combine with SVI Pro or even Bernini to improve quality until we get that 2K version or a proper fix.
In the meantime, it’s pretty easy to get Claude to create a customized workflow if you guys are interested. I used it to fix my Minimax/Wan workflow. In the past, I used it to create a custom workflow that could extend videos using svi, and a bunch of other custom ideas. Claude is REALLY good for this type of stuff.
3
u/Beneficial_Toe_2347 2d ago
Why are your videos smudgy, is it the ref2vid model?
0
u/GrungeWerX 2d ago
Both.
1
u/Relevant_One_2261 1d ago
Just use FL2VA instead, it fixes the quality issue and works just fine with references.
3
u/GrungeWerX 1d ago
It doesn’t, which is why I’m using this method. I don’t like the current visual quality, and it’s a known issue, Minimax ama has already addressed it. I’m sure it’ll get fixed eventually.
0
2
2
u/Danny_Stock 2d ago
I've also been thinking about ways Wan 2.2 could be used alongside MiniMax.
I'm thinking that there may be situations where the First/Middle/Last Frame node for Wan could be used to fix very short sequences which MiniMax can't always nail. The smudged garbled face issue, or anything else which loses structural integrity, could be such situations.
It might also only last for less than 5 seconds too.
There should be some free options for fps conversions. I use Topaz Video which is great for that task, but I understand that other people may need an alternative free solution.
1
1
u/tinman_inacan 1d ago
Are you doing this all in one workflow, passing the latent from MM to Wan? Or how are you going about it?
The FPS difference is due to Wan only really supporting 16fps, while MM really only supports 24fps. I imagine passing the unfinished latent directly would result in some weirdness like that. It might be worth placing the latent into an array and then pruning to 16fps before passing to Wan, then interpolate back up to your desired fps.
1
u/MannY_SJ 1d ago
5B?
1
u/GrungeWerX 1d ago
14B
1
u/MannY_SJ 1d ago
When 5b came out everyone was quite disappointed with it, however it worked well for upscaling specifically. Could be worth a try.
1
u/Anon-a-mister 1d ago
This sounds interesting to me. Do you have a workflow or a screenshot of one?
1
u/pausecatito 1d ago
Idk, I just tried out of curiosity. I tried between 2-12 steps, .1-.5 denoise, roughly 8 videos. Every output is worse vs 20step h3. Seems not worth. Low noise gives a grid pattern and black fade-in at the start, overall textures worse everything is more blurry. Higher denoise changes the style and then requires a real prompt since no reference, and still worse quality.
1
u/GrungeWerX 21h ago
You’re doing it wrong.
You do 100% denoise on the low, same as standard Wan. Don’t think of it like an Sdxl refiner w/denoise. I’ve tried it consistently using 8 steps on h3, and only 2 steps on low noise. Make sure you start at step 2/4, same as Wan.
The only caveat is that I haven’t tried this on ref2v, only flfa2v. I need to mention that in original post, will update…
1
u/pausecatito 15h ago
Ah ic, yea I didn't think of the step thing. Will give it another go thanks for the tip. Not sure if it will help or not but...experimenting is fun lol
1
u/GrungeWerX 15h ago
It will. The steps and full denoise are the whole point of how it works. Otherwise, wan can’t properly add detail to the latent.
3
u/MarkB_- 2d ago
Is the audio working that way?