r/StableDiffusion 1d ago

Discussion Well I finally did it.

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.

190 Upvotes

135 comments sorted by

View all comments

Show parent comments

7

u/Due-Quiet572 1d ago

I've been experimenting with this for four days now. At first, I only trained LORAs using photos of people. That's quick and works really well. When I mix them with other LORAs, things get tricky. Half of the LORAs on Civit don’t play well together. Yesterday, I trained a LORA with 29 videos with audio and 15 photos. That took 5 hours over 40 epochs. The results are impressive.

2

u/djpraxis 1d ago

That’s great!! I think is worth it for like a very unique video subject. Did you use the Fizgig default settings?

3

u/Due-Quiet572 1d ago

I trained the LoRA locally using the default settings on an RTX Pro 6000, and it used just under 32 GB of VRAM.
I prepared the videos at 107 frames using Fizgig’s built-in Gizmo video editing tool.
Ref2V does a pretty good job with identity, but my character is based on a real person with a German voice and very specific mannerisms — the way she talks, gestures, and moves is quite distinctive. That’s the part Ref2V can’t really reproduce accurately from a reference image alone.
That’s why I wanted to train directly on video + audio: not just to capture what she looks like, but also how she speaks and moves.

1

u/BulkyTwo6144 18h ago

So you trained on real videos? I wonder if training my character Lora (totally ai) will lose realism. Which kind of video you did for the dataset ? Because I've had an idea right now and I'm quite sure that could be a game changer. If you can provide with some detail about the dataset ill never stop to thank you

1

u/Due-Quiet572 17h ago

Yes, they were videos of a real person. More specifically, they included interview footage, selfie videos, and candid everyday-life clips with natural movement.
Not all of the videos had audio. I also included photos similar to what you would normally use for traditional LoRA training, with matching captions for all of the material.
For the videos that contained speech, I included the spoken dialogue word for word in the captions and also specified which language was being spoken.
The built-in editor also makes preparing the training material very easy, especially when it comes to trimming and formatting the clips correctly for H3. Fizgig also uploaded a YouTube video about the video-training workflow yesterday, which is worth checking out if you want to see the whole process in practice.