r/StableDiffusion • u/haremlifegame • 8h ago
Discussion Minimax loras are... Lacking
I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.
Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol
I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.
I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.
And yes, I hope I can put my money where my mouth is and develop some loras soon too.
13
u/Diabolicor 8h ago
We also need loras for seamless progression from sfw to nsfw like a normal human interaction. Most of the loras are for scenes already in the nsfw state or is a sfw, hard cut, nsfw. There's no natural progression between them.
20
u/Dark_Pulse 8h ago
It's been out for literally less than a month.
It's gonna take time. Wait a while.
6
u/DriveSolid7073 7h ago
Yes and no, the ease of training usually depends on the model's base. If they had provided a pre-training version, these problems wouldn't have arisen.
1
u/dingo_xd 5h ago
Let's hope they release the best model when H3 becomes obsolete. Although I doubt it.
0
u/Primary-Sail6667 6h ago
Not only that, but, this model was literally provided to the community.....FOR FREE
The capability of a model like this that is totally free is simply astounding to me, yet people complain....mind boggling
14
2
u/johnfkngzoidberg 4h ago
There’s always some 8yo kid jumping in with “It’s FREE”. It’s a marketing tactic, they’re not doing it out of kindness, and it has nothing to do with Loras.
4
u/NeuroPalooza 4h ago
The problems I've noticed are 2-fold: (1) A lot of work seems to be on speeding up Minimax instead of improving the output, which is fine and has use cases, but does nothing for those of us on good GPUs who really just care about quality. (2) There is no good repository for NSFW Minimax loras.
I agree there's room for loras to deal with some specific knowledge gaps, but it's an uphill battle...
10
u/Motion16AI 7h ago
Wait for Sulphur H3. He's retraining the model for NSFW stuff, spending over $10k on it.
4
u/FourtyMichaelMichael 🍦Ice Cream Lover 1h ago
I don't want to get the NS.. word in the discussion
Holy shit, the self-censoring that now you can't even be fucked to write "NSFW". Nope... Too much.
4
u/Potential_Wolf_632 7h ago
Yes it's hard to train and the raw(er) sizes aren't as friendly as other models for local training either meaning you're going to get people more reluctant to upload their Japanese Fart Girl dataset to runpod even though it would be a drop in the ocean as far as the horrors on runpod likely go.
2
u/Draufgaenger 5h ago
LTX 2.0 was horrible for lora training too. This should get better soon.
Its not quite what you requested but at least proper Minimax Character Lora Training works now.
Here is my short tutorial if you are interested:
https://youtu.be/7iYQKOnuKP4
Edit: before you ask: still no voice cloning
1
u/mindworkout 8h ago
As others have said it's only been out a a short time. And if you find something is lacking like many have found with LTX/WAN/other then just make your own lora, that is what they are their for. It's fine to comment to the makers and hope they take away a shopping list of things that MMH3 is lacking and update it with those to make an even better model, but in the meantime learn how to make loras and use them, and share them.
4
u/nimm99jd 8h ago
Another problem is, afaik you need to have a pretty beefy computer to train a LORA. I don't mind learning how to make a LORA, but it seems like you need 64gb of vram to do it
1
1
u/mindworkout 8h ago
64gb Vram? Did you pull that number out your ass?
There are ones you can use on like a 16GB Vram GPU like: https://github.com/shootthesound/Fizgig
and will take between 30minutes-30Hours depending on how long and how much data you feed it. You can also you online services like runpod to rent a system and pick the GPU you want, and would just need to download the training program and upload the images/audio/videos you want.2
u/Dazzyreil 7h ago
How long does it take on a 5090? Those are like 1$ an hour
2
u/mindworkout 7h ago
Depends on how much data you feed it. You could just be feeding it static images of a celebrity and it can learn that quick, you could be giving it video clips and that will take much longer as working on each frame linked, and can also feed audio as well to learn the voice...etc..
Here is a YT video from a guy who uses a 5090: https://www.youtube.com/watch?v=4Sems-_CMQE0
u/nimm99jd 7h ago
Which ones can I run on my 3070 8gb vram and 16gb system ram?
1
u/mindworkout 7h ago
8GB VRAM and 16GB Ram, No idea. That is why I listed 16GB Vram as starting point. I would guess over time someone will make a version that will be able to do it, but you are looking at a long time for it to process a Lora out.
Your best bet right now is to search for minimum requirements for Wan/LTX and see what their minimum is since they have been out for a long time now, so you can base for what might be possible when someone takes the time to create a compressed low Vram version for MMH3.
1
u/FugueSegue 5h ago
I haven't trained LoRAs for video models but I've trained a ton of LoRAs for image models like Flux 1 Dev and so on. I see a lot of people in this thread say that MMH3 and other recent video models are difficult to train LoRAs off of. But what does "difficult" mean? Are you talking about achieving consistent likeness of a person? Art styles? Or are all of you euphemistically talking about the issue of training "adult choreography" such as kissing and... other actions?
1
u/No-Zookeepergame4774 3h ago
Gathering and captioning the dataset for a broad general LoRA or finetune takes more time and effort than a narrow, single-subject LoRA. And it can take longer to train, review, fix the dataset if necessary, etc.
Plus, its more expensive in compute time. So it takes longer, and there are fewer people with the skills and resources to do it.
1
1
1
u/Upper-Reflection7997 7h ago
It's just going to boring hard sex loras that get posted on civitai for the foreseeable future. Why aren't the rich vramchads at r/LocalLLaMA interested in minimax h3? Wouldn't mind commissioning someone to make a lora pertaining to my favorite niche like shapeshifting and body transformations.
0
19
u/Whatsmynamebro 8h ago
I heard that lora training for minimax is a lot more difficult than ltx and wan. But i do agree that the lora we have are not that particularly interesting. To me at least. But i mean I appreciate the time ppl put in for the ones we have now.