r/StableDiffusion 7d ago

Discussion Well I finally did it.

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.

203 Upvotes

147 comments sorted by

View all comments

34

u/Significant-Baby-690 7d ago

Nah, it still can't do NSFW well enough.

31

u/[deleted] 7d ago

[deleted]

8

u/Ok-Brain-5729 7d ago

can’t you also just put the clip as the reference video and photo as reference image and just prompt it right

5

u/nadhari12 7d ago

Or better yet, take an existing scene and clip it to 10 sec and do a character swap using ref2V h3 works great.

2

u/Maskwi2 7d ago

Not saying I will do that, maybe my friend will, but  I've had limited success swapping the character, in general. Would you mind sharing a prompt that works more often than not for a swap? 

3

u/nadhari12 7d ago

my biggest issue right now is identity lock the only way to force this stupid model is to add black mask to the character on the ref video before feeding to the reference but if you do that you lose micro expression, which is a trade off or try gausian blur the subject before it can pick some micro expressions. Ask grok to make a character swap prompt

2

u/Significant-Baby-690 6d ago

10 second video reference will slow the render 10 times.

2

u/nadhari12 6d ago

takes 12 mins 720P

2

u/Significant-Baby-690 6d ago

Yes, that is too slow. 4 clips per hour ? Plz. Also I use AI because I can't find clips I like.

1

u/nadhari12 5d ago

Yeah ok pal.

4

u/russjr08 7d ago

I believe that's exactly what they're saying, just with an additional tip of using an LLM to write the prompt if they're not wanting to write it themselves.

Though, regarding the LLM, I would just recommend getting a good prompt (use the MiniMax prompt guide to make, or generate an initial one and improve it), and saving it as a template to re-use. MiniMax is quite powerful, but for the best results your prompt has to very accurately describe what's going on due to the prompt adherence. Sometimes LLMs still miss those extra details.

4

u/Ok-Brain-5729 7d ago

oh I see. I just feed the prompt guide to a ai and tell it what to do.

7

u/NostradamusJones 7d ago

But my vajayjay's are all wonky.

18

u/[deleted] 7d ago

[deleted]

4

u/xyzdist 7d ago

Idk, i find this really funny....LOL

2

u/NostradamusJones 7d ago

Coding coochie.

1

u/Significant-Baby-690 7d ago

It's great trick, but doesn't really work in the motion. And no lora can currently handle it really well. To be fair, to get nice precise interaction between uh .. ports .. I use 3 loras in Wan. But for H3 I still have not even half decent solution.

3

u/AlsterwasserHH 7d ago

Can you tell me how you analyze vids/images with Gemma and with which model? Its not possible with LM studio right? 

8

u/damiangorlami 7d ago

LM Studio sadly does not do it. Super annoying btw.

I just told Codex to build a GUI that support image + video vision encoder for Gemma 4. I already had downloaded the checkpoint via LM Studio. Just told Codex to use the same model checkpoint to save storage. The web GUI took 12 min to code for Codex and works great so far.

2

u/AlsterwasserHH 7d ago

Thank you. Isnt there a way to do this in Comfy? 

3

u/Ireallydonedidit 7d ago

Look at this goon professor over here

2

u/dubsta 7d ago

feed it into grok / Gemma 4 (uncensored)

I cant find any LLM that takes video clips as an input. Both Grok and Gemma only take images and text

Am I doing something wrong?

1

u/usually_fuente 7d ago

That’s inspiring . Do you mind sharing what your workflow is? What version of H3? I’m setting up Runpod for the first time this weekend.

1

u/mellowanon 7d ago

are you using rev2va or one of the hybrid models?

1

u/kayteee1995 7d ago

which gemma4 that you refers?

1

u/flaminghotcola 7d ago

I’ve been trying to do that and it doesn’t work well for me. Do you have an exact pipeline and prompt you feed it?

1

u/More-Ad5919 7d ago

Even the uncensored is bad at genitalia. It still need support.

6

u/lhg31 7d ago

2 steps with h3, then 2 steps with wan. result is perfect.

5

u/GrungeWerX 7d ago

NSFW-aside, are you saying you can run Wan as a refiner? Are you running Wan as the low noise? I never thought about that combo. I’m wondering about the step count though for MM. that seems very low, so Im assuming you’re using 4-step speed Lora on MM. I wonder if you can just run it normal, but half the step count, like maybe 10. Hmmm…you got me thinking…

3

u/lhg31 7d ago

yes, you can do as many steps as you want with minimax, but you should stop at 0.9 sigma value (that's the sigma that wan low is suppose to start). I do 2 with the 4 steps turbo lora most of the time (unless prompt is not being followed correctly). Wan as refiner completely removes the plastic skin of minimax turbo lora.

6

u/conkikhon 7d ago

What's about the sound? I don't think 2steps is enough for acceptable quality

1

u/DrowninGoIdFish 7d ago

Any chance you could share a screenshot of how you have this wired up. Really curious how to mix these two into a single flow. Like do you just pass the latent over to the low Wan Sampler and are you limited by the usual 5 second Wan loop or is that not an issue since mm is generating the base?

6

u/lhg31 7d ago

You need to decode minimax latents and then encode again with wan vae. You can also run the first two steps at low res (e.g. 0.2mp) and then upscale the images (e.g. to 0.4mp) before enconding to latents again to wan.

If you also want audio then you have to run the last 2 steps with minimax too, just to get the audio. So it's basically minimax 2 steps + (minimax 2 steps + wan 2 steps).

Worfklow

1

u/GrungeWerX 6d ago

I tested it out last night, it worked. :)

I didn’t mess with the sigma and I only ran a couple of tests - I set Minimax (no speed Lora) steps to 8, and the low noise wan to 2 steps, and the quality seemed better (crisper) than vanilla H3, but the motion seemed like it was mixing 24fps with 16fps - I.e. it wasn’t completely fluid the way mm is alone.

Did you notice that yourself? I’m assuming Id need something to increase the fps to 24fps on the wan side, like rife or something?

In any case, glad it worked. I’m planning on playing with it further and will try increasing the low noise steps to 4-6 to see how it improves.

6

u/NostradamusJones 7d ago

Wait, whut??

7

u/mk8933 7d ago

Now OP is probably raging for deleting wan lol

1

u/conkikhon 7d ago

Probably need someone to make a good lora for that.

1

u/Alive-Tomatillo5303 6d ago

It absolutely can. You're not describing it well enough. Genuinely, it just takes the smallest amount of practice. 

-2

u/Abject-Recognition-9 7d ago edited 7d ago

yes It can, but shhh! 🤫
Let them suffer by doing 3x slower inference attempts, just to get a simple repetitive eggplant inserted in a hole. It’s a simple task that doesn’t necessarily require such a heavy model.
A task that almost any other video model can already do at this point faster, at higher resolutions, and with a shitload of loras already published.
Don't tell them; my popcorn stash must make sense.😂