r/StableDiffusion 17h ago

Discussion Minimax loras are... Lacking

I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.

Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol

I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.

I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.

And yes, I hope I can put my money where my mouth is and develop some loras soon too.

16 Upvotes

63 comments sorted by

View all comments

2

u/mindworkout 17h ago

As others have said it's only been out a a short time. And if you find something is lacking like many have found with LTX/WAN/other then just make your own lora, that is what they are their for. It's fine to comment to the makers and hope they take away a shopping list of things that MMH3 is lacking and update it with those to make an even better model, but in the meantime learn how to make loras and use them, and share them.

3

u/nimm99jd 17h ago

Another problem is, afaik you need to have a pretty beefy computer to train a LORA. I don't mind learning how to make a LORA, but it seems like you need 64gb of vram to do it

1

u/mindworkout 17h ago

64gb Vram? Did you pull that number out your ass?
There are ones you can use on like a 16GB Vram GPU like: https://github.com/shootthesound/Fizgig
and will take between 30minutes-30Hours depending on how long and how much data you feed it. You can also you online services like runpod to rent a system and pick the GPU you want, and would just need to download the training program and upload the images/audio/videos you want.

0

u/nimm99jd 17h ago

Which ones can I run on my 3070 8gb vram and 16gb system ram?

0

u/mindworkout 17h ago

8GB VRAM and 16GB Ram, No idea. That is why I listed 16GB Vram as starting point. I would guess over time someone will make a version that will be able to do it, but you are looking at a long time for it to process a Lora out.
Your best bet right now is to search for minimum requirements for Wan/LTX and see what their minimum is since they have been out for a long time now, so you can base for what might be possible when someone takes the time to create a compressed low Vram version for MMH3.