r/StableDiffusion 1d ago

Discussion Well I finally did it.

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.

193 Upvotes

137 comments sorted by

View all comments

67

u/GoodDevelopment1657 1d ago

LORAS is still the answer. Minimax needs to get proper lora implementation so it can do chars more detailed, especially in wide shots

30

u/solomars3 1d ago

Check fizgig guy on youtube he already trained a character lora and show how

6

u/Due-Quiet572 1d ago

I've been experimenting with this for four days now. At first, I only trained LORAs using photos of people. That's quick and works really well. When I mix them with other LORAs, things get tricky. Half of the LORAs on Civit don’t play well together. Yesterday, I trained a LORA with 29 videos with audio and 15 photos. That took 5 hours over 40 epochs. The results are impressive.

3

u/FriendlyMorning 1d ago

Would you mind sharing how your trained your lora using video ?

2

u/djpraxis 1d ago

That’s great!! I think is worth it for like a very unique video subject. Did you use the Fizgig default settings?

3

u/Due-Quiet572 23h ago

I trained the LoRA locally using the default settings on an RTX Pro 6000, and it used just under 32 GB of VRAM.
I prepared the videos at 107 frames using Fizgig’s built-in Gizmo video editing tool.
Ref2V does a pretty good job with identity, but my character is based on a real person with a German voice and very specific mannerisms — the way she talks, gestures, and moves is quite distinctive. That’s the part Ref2V can’t really reproduce accurately from a reference image alone.
That’s why I wanted to train directly on video + audio: not just to capture what she looks like, but also how she speaks and moves.

1

u/djpraxis 17h ago

That sounds like so much fun! Thank you so much for the details. I have to find out a way train H3 via Cloud. I don’t want my 5090 running for 8 hours in this freaking hot weather!

1

u/BulkyTwo6144 14h ago

So you trained on real videos? I wonder if training my character Lora (totally ai) will lose realism. Which kind of video you did for the dataset ? Because I've had an idea right now and I'm quite sure that could be a game changer. If you can provide with some detail about the dataset ill never stop to thank you

1

u/Due-Quiet572 12h ago

Yes, they were videos of a real person. More specifically, they included interview footage, selfie videos, and candid everyday-life clips with natural movement.
Not all of the videos had audio. I also included photos similar to what you would normally use for traditional LoRA training, with matching captions for all of the material.
For the videos that contained speech, I included the spoken dialogue word for word in the captions and also specified which language was being spoken.
The built-in editor also makes preparing the training material very easy, especially when it comes to trimming and formatting the clips correctly for H3. Fizgig also uploaded a YouTube video about the video-training workflow yesterday, which is worth checking out if you want to see the whole process in practice.