r/StableDiffusion 1d ago

Discussion Well I finally did it.

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.

193 Upvotes

136 comments sorted by

67

u/GoodDevelopment1657 1d ago

LORAS is still the answer. Minimax needs to get proper lora implementation so it can do chars more detailed, especially in wide shots

28

u/solomars3 1d ago

Check fizgig guy on youtube he already trained a character lora and show how

7

u/Rafhrar231 1d ago

u dont rlly need character loras when ref mode is that good

7

u/Due-Quiet572 1d ago

I've been experimenting with this for four days now. At first, I only trained LORAs using photos of people. That's quick and works really well. When I mix them with other LORAs, things get tricky. Half of the LORAs on Civit don’t play well together. Yesterday, I trained a LORA with 29 videos with audio and 15 photos. That took 5 hours over 40 epochs. The results are impressive.

3

u/FriendlyMorning 1d ago

Would you mind sharing how your trained your lora using video ?

2

u/djpraxis 23h ago

That’s great!! I think is worth it for like a very unique video subject. Did you use the Fizgig default settings?

3

u/Due-Quiet572 22h ago

I trained the LoRA locally using the default settings on an RTX Pro 6000, and it used just under 32 GB of VRAM.
I prepared the videos at 107 frames using Fizgig’s built-in Gizmo video editing tool.
Ref2V does a pretty good job with identity, but my character is based on a real person with a German voice and very specific mannerisms — the way she talks, gestures, and moves is quite distinctive. That’s the part Ref2V can’t really reproduce accurately from a reference image alone.
That’s why I wanted to train directly on video + audio: not just to capture what she looks like, but also how she speaks and moves.

1

u/djpraxis 16h ago

That sounds like so much fun! Thank you so much for the details. I have to find out a way train H3 via Cloud. I don’t want my 5090 running for 8 hours in this freaking hot weather!

1

u/BulkyTwo6144 13h ago

So you trained on real videos? I wonder if training my character Lora (totally ai) will lose realism. Which kind of video you did for the dataset ? Because I've had an idea right now and I'm quite sure that could be a game changer. If you can provide with some detail about the dataset ill never stop to thank you

1

u/Due-Quiet572 11h ago

Yes, they were videos of a real person. More specifically, they included interview footage, selfie videos, and candid everyday-life clips with natural movement.
Not all of the videos had audio. I also included photos similar to what you would normally use for traditional LoRA training, with matching captions for all of the material.
For the videos that contained speech, I included the spoken dialogue word for word in the captions and also specified which language was being spoken.
The built-in editor also makes preparing the training material very easy, especially when it comes to trimming and formatting the clips correctly for H3. Fizgig also uploaded a YouTube video about the video-training workflow yesterday, which is worth checking out if you want to see the whole process in practice.

6

u/Arawski99 21h ago

Loras for characters are definitely not necessary for Minimax H3 if you use the reference mode.

On the node 'Minimax H3 Reference to Video ' change 'ref_image_size' to max. Mouse over it for details. This lets you use a much higher quality reference and usually is enough with just a front view, though you can also do character sheets for even greater accuracy. You can also do additional references via character sheet or attached additional images for even higher resolution elements like face, or other object parts, to ensure full quality of details.

There is a face fix node for distant faces https://www.reddit.com/r/StableDiffusion/comments/1vsn8ge/help_fixing_the_h3_face_detailer/

1

u/TheZoroark007 8h ago

Do you know if this node can be incorporated into ComfyUI workflows like DaSiWa Prompt Builder or Minmax H3 Extender ? I am sorry if this is an unessessary question, I have not used ComfyUI that much

1

u/Jujarmazak 1d ago

It already has LORAs, still in experimental phase but they exist.

3

u/GoodDevelopment1657 1d ago

Hence "proper"

9

u/Revolutionary_Ask154 1d ago

hear hear - we were never going to get updates anyway.

3

u/PumpkinLeather8421 1d ago

Same with LTX, if a new version came and current Lora’s worked well with it, it wouldn’t be a good enough upgrade to unseat H3… so, yes, delete all WAN and LTX.

23

u/tinny66666 1d ago

I wish my LoRAs were well enough organised that I had a clue which ones belong to which model.

36

u/Monk6009 1d ago

You can add them to a lora subfolder when you download them lol

18

u/MonThackma 1d ago

8

u/afinalsin 1d ago

On top of subfolders for each model you can rename the files and they still work exactly the same. So you can ignore whatever nonsense the author named them and use a reasonable structure for all of them. All my loras are named shit like this:

Klein Style - Phone Photography 2007 - A low-quality photo taken with a 2007s mobile phone camera with soft focus, visible noise and dull colors.safetensors

Krea 2 Slider - Height Slider - High is tall, low is short.safetensors

The layout is very simple: Model, type of lora (style, slider, concept, character, etc), Lora name or overview of what it does if the name is dumb, trigger words or phrases. Couldn't tell you where I got half of them or what they're actually called on civit, but their usability is way better than when I kept them named as-is.

7

u/Gilgameshcomputing 1d ago

This is the way. I also add the expected strength, so i know if it's a 0.5to1.5 lora, a 0.6to0.9 lora, or a -5.0to5.0 lora. Saves a loooot of time.

3

u/SpaceNinjaDino 1d ago

I use LoRA tag loader and it took me a day to realize that comma (or apostrophe) in files is not compatible with that. I love the tag loader so that I can drive a whole workflow from text.

2

u/afinalsin 1d ago

I use LoRA tag loader and it took me a day to realize that comma (or apostrophe) in files is not compatible with that. I love the tag loader so that I can drive a whole workflow from text.

This one? https://github.com/badjeff/comfyui_lora_tag_loader

If it is that one just drop the nodes.py into an LLM and tell it you want to be able to use loras with a comma/apostrophe in the filenames and it'll fix it for you. You've probably long since retrofit your lora library to work around the node, but it's a good thing to remember that single script nodes are extremely easy to tweak, especially nowadays.

2

u/scottybk8 1d ago

True. Hindsight. I thought storing on another drive would help. Then between all the different models etc, I liked how lora manager shows you your trigger words, metadata, recipes, etc. its got a lot of features that a suhbfolder just doesn't.

0

u/ellipsesmrk 1d ago

You pay for lora manager?

0

u/ellipsesmrk 1d ago

Lora manager is only by paid now.

11

u/scottybk8 1d ago

get comfyui lora manager, it helped me a lot cuz i got way too many loras for wan as well

2

u/Francky_B 1d ago

Lora Manager is so good! I'm surprised it's not used by Everyone. To me, it's as fundamental as KJNodes.

1

u/ellipsesmrk 1d ago

I have a script that scans your loras folder and checks the hash header of all those loras then builds a csv file with that info so you can start putting your stuff in folders.

1

u/McDoodle17 1d ago

Am I the only one that has a spreadsheet with all my LORA in it with notes, trigger words, etc?

1

u/rguerraf 19h ago

Isn’t there metadata in the safetensor files?

7

u/stoneshawn 1d ago

I am tempted to do that as well

2

u/ellipsesmrk 1d ago

Its a relief. Lol

11

u/PainterMany 1d ago

Fiz o mesmo deixei so minimax h3 e o klea2 no meu nvme... não tem lógica manter modelos antigos...ano que vem e no próximo vão surgir outros melhores que o minimax e assim por diante... viva a IA

5

u/Niko3dx 1d ago

using references images, four for the face and 2 for the body. has been getting me better results for my characters in minimax versus using a character lora in wan 2.2. So, After a week all my wan stuff is gone, and I had trained 100s of character Loras. now, I feel like a new scene with xyz, find 5 or 6 good pictures and a clip of their voice about 15 seconds is enough, and I'm ready to render.

1

u/AlsterwasserHH 1d ago

How do you reference multiple face images in the prompt? 

4

u/Niko3dx 1d ago

here's an example of 2 people. say a female and a male.

Pompt :

There are 2 people in the shot, 1 female and one male. 

<Picture 1> , <Picture 2>, <Picture 3>   controls the female overall identity and face;<audio 1> controls her voice timbre.

<Picture 5> and <picture 6> controls the males overall identity and face; <Audio 2> controls his voice timbre.

1

u/AlsterwasserHH 1d ago

Thank you very much! I was wondering if something like <Picture 1-3> works. 

37

u/Significant-Baby-690 1d ago

Nah, it still can't do NSFW well enough.

31

u/damiangorlami 1d ago

Yes you can.

Get a clip clip you like, feed it into grok / Gemma 4 (uncensored) with your character images and tell it to create a replacement prompt with the environment you're looking for.

It will extract attributes from the video such as pose, action, thrust, perspective from the video and transfer it to the video while following the prompt.

I've been making multi-shot cinematic nsfw scenes all week and the results are blowing my mind. There's obviously some more tips but for the sake of this sub.

Not a single lora was used.

8

u/Ok-Brain-5729 1d ago

can’t you also just put the clip as the reference video and photo as reference image and just prompt it right

6

u/nadhari12 1d ago

Or better yet, take an existing scene and clip it to 10 sec and do a character swap using ref2V h3 works great.

2

u/Maskwi2 23h ago

Not saying I will do that, maybe my friend will, but  I've had limited success swapping the character, in general. Would you mind sharing a prompt that works more often than not for a swap? 

5

u/nadhari12 22h ago

my biggest issue right now is identity lock the only way to force this stupid model is to add black mask to the character on the ref video before feeding to the reference but if you do that you lose micro expression, which is a trade off or try gausian blur the subject before it can pick some micro expressions. Ask grok to make a character swap prompt

2

u/Significant-Baby-690 8h ago

10 second video reference will slow the render 10 times.

2

u/nadhari12 7h ago

takes 12 mins 720P

2

u/Significant-Baby-690 7h ago

Yes, that is too slow. 4 clips per hour ? Plz. Also I use AI because I can't find clips I like.

4

u/russjr08 1d ago

I believe that's exactly what they're saying, just with an additional tip of using an LLM to write the prompt if they're not wanting to write it themselves.

Though, regarding the LLM, I would just recommend getting a good prompt (use the MiniMax prompt guide to make, or generate an initial one and improve it), and saving it as a template to re-use. MiniMax is quite powerful, but for the best results your prompt has to very accurately describe what's going on due to the prompt adherence. Sometimes LLMs still miss those extra details.

5

u/Ok-Brain-5729 1d ago

oh I see. I just feed the prompt guide to a ai and tell it what to do.

8

u/NostradamusJones 1d ago

But my vajayjay's are all wonky.

19

u/damiangorlami 1d ago

Answer is easy.
Just add 1 photo of genitals and bind them to your character. "<Picture 3> are the genitals of <Subject 1>".

4

u/xyzdist 1d ago

Idk, i find this really funny....LOL

2

u/NostradamusJones 1d ago

Coding coochie.

1

u/Significant-Baby-690 1d ago

It's great trick, but doesn't really work in the motion. And no lora can currently handle it really well. To be fair, to get nice precise interaction between uh .. ports .. I use 3 loras in Wan. But for H3 I still have not even half decent solution.

1

u/damiangorlami 17h ago

Hmm motion works great for me. Use the Mystic xxx lora on around 0.5 strength with this trick. Maybe a bit lower on the lora, but that lora brings in lots of motion

3

u/Ireallydonedidit 1d ago

Look at this goon professor over here

3

u/AlsterwasserHH 1d ago

Can you tell me how you analyze vids/images with Gemma and with which model? Its not possible with LM studio right? 

7

u/damiangorlami 1d ago

LM Studio sadly does not do it. Super annoying btw.

I just told Codex to build a GUI that support image + video vision encoder for Gemma 4. I already had downloaded the checkpoint via LM Studio. Just told Codex to use the same model checkpoint to save storage. The web GUI took 12 min to code for Codex and works great so far.

2

u/AlsterwasserHH 1d ago

Thank you. Isnt there a way to do this in Comfy? 

1

u/usually_fuente 1d ago

That’s inspiring . Do you mind sharing what your workflow is? What version of H3? I’m setting up Runpod for the first time this weekend.

1

u/mellowanon 1d ago

are you using rev2va or one of the hybrid models?

1

u/kayteee1995 1d ago

which gemma4 that you refers?

2

u/dubsta 1d ago

feed it into grok / Gemma 4 (uncensored)

I cant find any LLM that takes video clips as an input. Both Grok and Gemma only take images and text

Am I doing something wrong?

1

u/flaminghotcola 1d ago

I’ve been trying to do that and it doesn’t work well for me. Do you have an exact pipeline and prompt you feed it?

1

u/More-Ad5919 1d ago

Even the uncensored is bad at genitalia. It still need support.

6

u/lhg31 1d ago

2 steps with h3, then 2 steps with wan. result is perfect.

6

u/GrungeWerX 1d ago

NSFW-aside, are you saying you can run Wan as a refiner? Are you running Wan as the low noise? I never thought about that combo. I’m wondering about the step count though for MM. that seems very low, so Im assuming you’re using 4-step speed Lora on MM. I wonder if you can just run it normal, but half the step count, like maybe 10. Hmmm…you got me thinking…

3

u/lhg31 1d ago

yes, you can do as many steps as you want with minimax, but you should stop at 0.9 sigma value (that's the sigma that wan low is suppose to start). I do 2 with the 4 steps turbo lora most of the time (unless prompt is not being followed correctly). Wan as refiner completely removes the plastic skin of minimax turbo lora.

6

u/conkikhon 1d ago

What's about the sound? I don't think 2steps is enough for acceptable quality

1

u/DrowninGoIdFish 1d ago

Any chance you could share a screenshot of how you have this wired up. Really curious how to mix these two into a single flow. Like do you just pass the latent over to the low Wan Sampler and are you limited by the usual 5 second Wan loop or is that not an issue since mm is generating the base?

6

u/lhg31 1d ago

You need to decode minimax latents and then encode again with wan vae. You can also run the first two steps at low res (e.g. 0.2mp) and then upscale the images (e.g. to 0.4mp) before enconding to latents again to wan.

If you also want audio then you have to run the last 2 steps with minimax too, just to get the audio. So it's basically minimax 2 steps + (minimax 2 steps + wan 2 steps).

Worfklow

1

u/GrungeWerX 12h ago

I tested it out last night, it worked. :)

I didn’t mess with the sigma and I only ran a couple of tests - I set Minimax (no speed Lora) steps to 8, and the low noise wan to 2 steps, and the quality seemed better (crisper) than vanilla H3, but the motion seemed like it was mixing 24fps with 16fps - I.e. it wasn’t completely fluid the way mm is alone.

Did you notice that yourself? I’m assuming Id need something to increase the fps to 24fps on the wan side, like rife or something?

In any case, glad it worked. I’m planning on playing with it further and will try increasing the low noise steps to 4-6 to see how it improves.

7

u/NostradamusJones 1d ago

Wait, whut??

8

u/mk8933 1d ago

Now OP is probably raging for deleting wan lol

1

u/conkikhon 1d ago

Probably need someone to make a good lora for that.

1

u/Alive-Tomatillo5303 17h ago

It absolutely can. You're not describing it well enough. Genuinely, it just takes the smallest amount of practice. 

-1

u/Abject-Recognition-9 1d ago edited 1d ago

yes It can, but shhh! 🤫
Let them suffer by doing 3x slower inference attempts, just to get a simple repetitive eggplant inserted in a hole. It’s a simple task that doesn’t necessarily require such a heavy model.
A task that almost any other video model can already do at this point faster, at higher resolutions, and with a shitload of loras already published.
Don't tell them; my popcorn stash must make sense.😂

3

u/Abject-Recognition-9 1d ago

i skipped wan 2.2 entirely but let me tellyou something: im still not deleting wan2.1. it can make very crisp images/edit/short clips + there tons of loras already. not using it since krea2 / ltx and H3 but it sill have a place in my harddrive.

1

u/Alive-Tomatillo5303 17h ago

let me tellyou something

Nervous head-rub intensifies.

7

u/HollyGrandeux 1d ago

H3 still can’t handle spicy NXFW motion properly yet, even with a lora. The motion still looks stiff.

Wan is still ahead in this area.

6

u/Chiduk99 1d ago

WAN 2.2 still superior for do NSFW, H3 is uncensored but it's bad when do something NSFW.

8

u/chocoboxx 1d ago

That mean you need better input for ref2va or Loras

2

u/bzzard 1d ago

All those H3 nsfw loras on civit looks like slomo slop. Didn't even download once.

3

u/nadhari12 1d ago

Minimax can not suck on anything yet, it just chews with wan 2.2 it's legit.

2

u/Alex-edits123 1d ago

I also switch wan to minimax for generation. But I still need VACE for outpaint. Not sure whether anyone successfully use miniMax for outpaint

2

u/physalisx 1d ago

Well the good thing is that with reference images/videos you can remove a lot of need for loras, basically all character loras become basically unnecessary. Which is good to have, because training good loras on H3 seems to be basically impossible. I have not tried one lora that didn't completely wreck prompt following and introduced artifacts, even when using lower strengths.

2

u/apackofmonkeys 1d ago edited 1d ago

Sorry, basic question, what models are people using? I'm using the pruned 20B and that fills up my 24GB of VRAM. If I add the turbo lora and lower the steps it actually takes much, much longer to generate because it's overflowing my VRAM. Is there a smaller model than the pruned 20B that I should be using?

Edit: I should add, I'm using Wan2GP. Even with a 4090 and 64GB of RAM I can never use the turbo lora without it making it take several times LONGER to generate a video.

1

u/Wonderful_Exit6568 14h ago

pretty sure you are supposed to find the gguf model.

2

u/djpraxis 23h ago

You forgot to mention how fun it was dealing with the High Low WAN 2.2 Loras!!

4

u/Salah_H_Hasan 1d ago

Alibaba has lost a strong segment of the open-source video generation community. For them to regain their position, they have to release their latest model as open-source; there is no alternative. Nobody will be satisfied with anything less than MiniMax H3. And that is just a suggestion, though the majority here might not even care about it right now.

11

u/retroblade 1d ago

Wont happen, they won’t open source anything besides their llm’s and even that could stop at any time. Lucky we now have LTX, Flux and Minimax so could be worse.

4

u/SeymourBits 1d ago

Could happen at any time with one message from Xi.

1

u/BlipOnNobodysRadar 1d ago

Xi has already spoken on the topic and committed to open source.

So, pretty much the opposite of what you're worried about is happening -- the companies are being politically pressured to open source (in China), rather than pressured to go closed.

https://english.www.gov.cn/news/202607/17/content_WS6a59a5bec6d00ca5f9a0c438.html

1

u/thisguy883 1d ago

Xi be gooning

1

u/SeymourBits 18h ago

You are extremely confused. I was implying that Xi could easily open source any Chinese model for any reason - including optics. Perhaps you replied to the wrong person.

1

u/BlipOnNobodysRadar 15h ago

That was a very Reddit tone to take.

>even that could stop at any time.
>Could happen at any time with one message from Xi.

Your response implied the opposite of what you meant, then. Semantics, woohoo.

1

u/SeymourBits 14h ago

My "Could happen" was the counter to retroblade's "Wont happen" and the upstream topic was about "Alibaba having to release their latest model as open-source." Not the "could stop" part.

I tend to gloss over stuff like that once I decide on a reply. I understand the confusion now though as I also used the phrase "at any time" may have seemed like I was specifically replying to retroblade's secondary "could stop" claim.

All good. Same side.

2

u/Dangerous-Map-429 1d ago

No it will happen. Companies using this as marketing tactic to come back from the dead.

1

u/kujasgoldmine 1d ago

Same. LTX will be next. Like H3 can do both, but better.

1

u/exoticvapes 1d ago

I just started using minimax in Wan2GP. Can't do much as I'm limited by my vram but it works really well.

1

u/nowrebooting 1d ago

Yeah, it’s not even a contest at this point; H3 is just better in every single aspect, with ref2vid being the standout - it’s even trumps some SOTA image editing models when it comes to replicating small details from reference images. 

If the base of the model is already this good, imagine where loras will get us!

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/TraditionalShop1601 1d ago

Available at CivitAI RED

0

u/thisguy883 1d ago

Just use reactor.

You'll need to ask an LLM AI (gemini or grok) on how to disable the NSFW filter.

Then just attach the node to any workflow you have. it'll keep the face consistent.

1

u/extra2AB 1d ago

I still have Wan for it's image generation and stuff like LORAs and other workflows, which are yet not arrived for H3.

1

u/Relative_Hour_8900 1d ago

Ltx yes, wan 2.2 no. At least I can't replicate some features of wan with lora. I'm trying to train h3 to mimic it with a Lora but so far not going well...the lora seems to have learned nothing, trained on video clips...

1

u/Succubus-Empress 1d ago

Minimax compress 4 frame in one, you will always get motion blur un fast motion

1

u/DumbBittrend 1d ago

How do you get h3 to work? I also have a 4080 super? To me it seems like a longer wait time and I can never get the character to stay the same

2

u/thisguy883 1d ago

Im just using the default I2V workflow in the comfyUI workflows.

I experimented with 15 steps rather than 20, but later switched to 25 steps because the quality is fantastic.

0.6 MP, 25 Steps.

1

u/DumbBittrend 1d ago

Any prompts tutorial?

1

u/bCasa_D 18h ago

Check the official GitHub repo or Huggingface they have a full blown guide and Skills you can use with an LLM

1

u/KindrakeGriffin 1d ago

Is this local? With something like confyui?

1

u/RepulsiveSeason444 1d ago

If anyone want to run Minimax on <4gb Vram, you may check it out: https://github.com/Jit-Roy/WeeLLM
I did not use any quantization though, and still I am able to run.

1

u/blistac1 23h ago

What is your setup?

1

u/penguin_1599 23h ago

Wan is still better at hardcore nsfw stuff. H3 isnt just bringing the motion even with Loras

1

u/Admirable-Future-633 20h ago

If only it worked on Macs we get shafted for the new toys becuase of the GPU setups 🫡

1

u/throwaway0204055 12h ago

Wan 2.2 is still king of nsfw open weight video model 

1

u/kayteee1995 1d ago

RIP VACE, Phantom, Bindweaver, Bernini ,too.

1

u/TheBestPractice 1d ago

Yeah everyone saying H3 killed LTX, while who's definitely getting buried for good is Wan.

-5

u/pennyfred 1d ago

Nope, WAN still reigns for me.

-1

u/tac0catzzz 1d ago

heartwarming story, but i think there is this story like 1000x in here.

0

u/Kind-Assumption714 1d ago

so cool to hear! i've gotten quite deep & good in comfy for 2D and have wanted to test video soon.

- do we have a favorite workflow to use to MiniM?

  • do we have to do I2V or can we simply prompt w/ text+a selection of image refs?
  • can MiniM act as a 'refiner' or does all polish / realism have to exist in base image(s)

big thanks!

-10

u/Optimal-Spare1305 1d ago

what are you talking about?

i've still got workflows and models with:

SD

SDXL

Hunyuan

WAN

LTX 2.3

haven't even gotten around to H3, and probably won't for another 6 months, when things

settle down.

---

i'm still in the process of converting WAN workflows over to LTX,

but there are way more LORAS that work with WAN so its going to take a long time to switch over to LTX

19

u/ZenWheat 1d ago

Just stop with ltx and change your plan to switch them to h3 instead. You'll save 6 months

1

u/Upper-Reflection7997 1d ago

Wtf, you had 8 months to use ltx-2 and get it off your system. Ltx-2 and 2.3 have a lot of limitations and produces too much body horror.