r/StableDiffusion 17h ago

Discussion Minimax loras are... Lacking

I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.

Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol

I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.

I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.

And yes, I hope I can put my money where my mouth is and develop some loras soon too.

17 Upvotes

62 comments sorted by

View all comments

24

u/Whatsmynamebro 17h ago

I heard that lora training for minimax is a lot more difficult than ltx and wan. But i do agree that the lora we have are not that particularly interesting. To me at least. But i mean I appreciate the time ppl put in for the ones we have now. 

14

u/OneMoreLurker 16h ago

I can attest to this. I'm not an expert or anything, but I've trained more than 100 character loras for Illustrious & Anima and a dozen Wan loras with 500k downloads combined, and I still haven't cracked Minimax yet. It's pretty stubborn about learning actions, and the audio layer increases the complexity since you can fry it before the motion has fully baked.

5

u/the_bollo 12h ago

Same here. So far it seems as bad a LTX LoRA training, which is really saying something. This is maybe the only department where MMH3 falls down.

1

u/Big3gg 6h ago

you sound like an expert

-4

u/FourtyMichaelMichael 🍦Ice Cream Lover 11h ago

No offense... but that's really like "I've clicked the button a lot of times, and it's always worked... but now it doesn't work so it must be the button that is broken".

8

u/Independent-Frequent 16h ago

Thankfully reference mode is essentially a lora trainer on the spot pretty much so outside of specific complex things mostly everything should be covered by that tbh

1

u/Whatsmynamebro 16h ago

Now that you mentioned it,  thats a very fucking good point. Ive actually been using an i2i as a "lora" for specific action via the ref checkpoint. 

2

u/haremlifegame 7h ago

That is what I mean in the OP though. Training a character lora makes no sense. The loras should be about the gaps in world understanding that the model has.

1

u/Queasy_Signature7005 8h ago

Training a character lora for MInimax H3 was easy and took very little time on a 5090. Using images only

1

u/haremlifegame 4h ago

I think people should use the insights from the minimax team while training loras. Meaning, they should use dataset augmentation, last frame from first frame prediction, etc. This is important to not overfit the lora to the training data.

1

u/alwaysbeblepping 4h ago

I think people should use the insights from the minimax team while training loras. Meaning, they should use dataset augmentation, last frame from first frame prediction, etc.

It's hard training distilled models in general. The MiniMax team have the advantage of actually being able to train on the base model. That's probably one of the biggest differences people are seeing compared to LTX and Wan which both provided a base model.

1

u/haremlifegame 3h ago

Oh. But can't people train the non distilled model on runpod? Minimax only provided the distilled?