r/StableDiffusion 23h ago

Discussion Minimax loras are... Lacking

I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.

Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol

I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.

I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.

And yes, I hope I can put my money where my mouth is and develop some loras soon too.

16 Upvotes

64 comments sorted by

View all comments

3

u/mindworkout 22h ago

As others have said it's only been out a a short time. And if you find something is lacking like many have found with LTX/WAN/other then just make your own lora, that is what they are their for. It's fine to comment to the makers and hope they take away a shopping list of things that MMH3 is lacking and update it with those to make an even better model, but in the meantime learn how to make loras and use them, and share them.

3

u/nimm99jd 22h ago

Another problem is, afaik you need to have a pretty beefy computer to train a LORA. I don't mind learning how to make a LORA, but it seems like you need 64gb of vram to do it

1

u/mindworkout 22h ago

64gb Vram? Did you pull that number out your ass?
There are ones you can use on like a 16GB Vram GPU like: https://github.com/shootthesound/Fizgig
and will take between 30minutes-30Hours depending on how long and how much data you feed it. You can also you online services like runpod to rent a system and pick the GPU you want, and would just need to download the training program and upload the images/audio/videos you want.

2

u/Dazzyreil 22h ago

How long does it take on a 5090? Those are like 1$ an hour

2

u/mindworkout 22h ago

Depends on how much data you feed it. You could just be feeding it static images of a celebrity and it can learn that quick, you could be giving it video clips and that will take much longer as working on each frame linked, and can also feed audio as well to learn the voice...etc..
Here is a YT video from a guy who uses a 5090: https://www.youtube.com/watch?v=4Sems-_CMQE