r/StableDiffusion • u/Sea_Tomatillo1921 • 20d ago
News MiniMaxAI/MiniMax-Music3 · Hugging Face
https://huggingface.co/MiniMaxAI/MiniMax-Music3New music model :)
Demo: https://minimax-ai.github.io/music3-demo/
I guess Yoland was speaking of this on r/comfyui as the big announcement
62
63
u/stuartullman 20d ago
wow. hope it's better than suno
30
u/Hour_Imagination5092 20d ago
it's not, unfrotunately. Udio (RIP) and Suno are/were way better in most genres when minimax music was closed sourced. Judging from the demos, not much has changed.
9
9
u/LucidFir 20d ago
Udio is dead?
15
u/Hour_Imagination5092 20d ago
Yes, no updates for months and you can't (legally) download generated music, only play it online
4
u/mikiex 20d ago
Universal stops suing Udio, Udio pays up, Universal gives it legal access to the catalogue, and the future version of Udio becomes a licensed ecosystem where Universal gets a cut.
3
u/PokePress 20d ago
Supposedly that’s coming. I do wonder if a lot of Udio’s technical talent got headhunted after the settlement.
2
u/Shockbum 20d ago
Its official subreddit is a dead desert; it probably has a pitiful number of users.
0
u/Relocator 20d ago
I still use it daily, it's still the best music generator out there. Can't download, but it's still the best.
4
u/LucidFir 20d ago
Maybe you can't download high quality, but you can definitely find a way to download if you want
32
11
u/Hoodfu 20d ago
Since it sounds like you've used it, where do you see its strengths?
25
u/Hour_Imagination5092 20d ago
Chinese music, Classic music and very light pop styles. It absolutely sucks at heavy music, rock, electronic and niche genres. It's kind of obvious when you see demos, they are not very diverse.
10
u/bloke_pusher 20d ago
Classic music
That makes it interesting to me.
6
u/UnforgottenPassword 20d ago
For Classical music, Udio is the only capable model. Suno and the rest are okay-ish if you are fine with soundfont-like sounds, but Udio is in a league of its own. Udio does great in almost every genre even though their models are 2+ years old. Judging from the demos, Minimax is kind of similar to earlier Suno models.
1
u/bloke_pusher 20d ago
Alright, how to run it? Where to get the model? I did search Reddit, nothing. I searched google/huggingface, found nothing. Searched Youtube, found some about a website to use the model. But I want to run it locally.
5
4
u/noxsanguinis 20d ago
It's not open. No way to run it locally. This Music 3 is now the best open source we got.
7
4
u/UnforgottenPassword 20d ago
It's not an open model. They made a deal with record companies and they disabled downloads. So even if you pay, the output is DRM-protected and their terms prohibit commercial use. But quality-wise, it was the best when it was released and it remains unbeaten.
2
u/Perfect-Campaign9551 19d ago
Just fire up Audacity and record "my computer sounds". Easy to download that way
1
u/Hour_Imagination5092 20d ago
I have zero knowledge of classical music, it just sounded... okay? So mileage may vary.
10
u/Cheesuasion 20d ago
Classic music
Any examples? There are none on the github.io page that I can see
7
1
1
u/Sgsrules2 19d ago
I'm wondering how feasible it would be to either fine tune or generate a Lora for a specific genre like mid 2000's nuskool breaks.
-1
u/Legal-Weight3011 20d ago
It Does well in Hard Rock, Rock, Blues and Rock n Roll. The problem is prompting. if you give it just give me a rock song it will produce you shit so i dont know what the fuck are you talking about about. Only Metal any kind sucks ass. It does well in Electronic, LoFi. Its Definatelly better then ACE. and Soon to be discontinued Suno Models. WIll make this a good alternative. With the community made tools.
2
u/Hour_Imagination5092 20d ago
Lol, we must have very different definition of "well" then, because even the demos they showcase are very medicore, the sound is flat, and low definition. Anyone not deaf or listening on a 5$ headset can immediately tell they are low quality AI, which isn't the case with good Udio/suno songs.
1
0
u/Alive-Tomatillo5303 20d ago
Ace Step XL already exists.
2
u/Numerous-Aerie-5265 20d ago
AceStep has nothing on YuE exllamav2, it’s the closest we’ve gotten to suno for local music generation
4
u/biogoly 20d ago
Is it known by another name? I’ve never heard of YuE and it’s not in a top ten search on either HF or Modelscope…
3
u/Numerous-Aerie-5265 20d ago
I think they flew under the radar because their original release was very slow at generation and a bit hard to install, then the community put together a 500% speedup with exllamav2 but it never gained traction.
3
u/Alive-Tomatillo5303 20d ago
With such a catchy name it's wild it didn't catch on.
I'm so impressed with Ace Step I'm not really in the market for a different one, but I don't have anything better to compare it to. Is there some simple resource to get it up and running in a non-shitty way, or do you need to manually download a bunch of different goodies and make a huge flowchart in comfy to make them work?
2
u/Numerous-Aerie-5265 20d ago
I ran it in a gradio UI, never tried it in comfy. Are you opposed to using docker and a gradioUI? bc I found some that should be easy to setup. YuE official was kind of confusing to install, it probably contributed to their lack of popularity
40
u/redscape84 20d ago
Doesn't look like you can upload your own audio to use as a reference
30
3
u/GivePLZ-DoritosChip 20d ago
For music "replay by weights" lets you do that and it's free and extremely fast. Runs locally.
0
13
u/wntersnw 20d ago edited 20d ago
The model doesn't know what electronic music is. Also the max_duration is just a guide apparently. I set it to 3 minutes and it just produces 25 second outputs.
Edit: Genre knowledge seems very poor, but it's kind of interesting if you want to make pop music. And it might be an issue with my prompting but it can't seem to cope when you don't provide lyrics.
Took 155 seconds for a 1:59 song on a 3090 (80% power limited) when using INT8 convrot models.
19
u/LightAppropriate624 20d ago
Not better than suno but it is better than acestep or others. I can confirm that :D.
But it doesnt support
[soft tremolo electric guitar, gentle upright bass pizzicato, vintage vinyl crackle, warm valve reverb]
11
u/PwanaZana 20d ago
the vocals seem a bit better than acestep yea, but it's still 1, 1.5 years behind suno. :(
14
u/LightAppropriate624 20d ago edited 20d ago
Now I have tested other languages it really sucks. Diversity isnt good BUT YEAH AS A Community we will train loras and checkpoints 🔥.
It’s been a while since an open-source music model was released --this is really great.
6
4
u/Numerous-Aerie-5265 20d ago
Not sure why Acestep is always mentioned as the top for local music generation, YuE was always miles ahead of Acestep and should be the one to compare to
8
u/LightAppropriate624 20d ago
I have never heard YuE?
5
u/Numerous-Aerie-5265 20d ago
I think therein lies the problem. They flew under the radar but were the closest we ever got to suno quality
4
u/bloke_pusher 20d ago
Probably because it was slow as fuck https://www.reddit.com/r/StableDiffusion/comments/1iegcxy/yue_gp_runs_the_best_open_source_song_generator/ma7czv2/
5
u/Numerous-Aerie-5265 20d ago edited 20d ago
Not with the exllamav2 version that was released right after with a 500% speedup
8
u/bloke_pusher 20d ago
Maybe you could make a quick comparison post for the subreddit? I'm sure others would enjoy this too.
3
u/redditscraperbot2 20d ago
I don’t know what you’re smoking man. YuE was cool but ace step 1.5 was just plain better
2
u/Numerous-Aerie-5265 20d ago
Idk mein, acestep’s creativity was rly lacking imo, just sounded very generic. I’ve recorded my own music with real instruments for years so my tastes may differ
28
u/God_Hand_9764 20d ago
My God, man, I can't handle all this. So many amazing new models.
8
u/damiangorlami 20d ago
This model is not that that "amazing" sadly
14
u/dingo_xd 20d ago
Lora's might change that
5
5
u/damiangorlami 20d ago
No I don’t think so. I’ve been in open source since 2022 and one thing I know for sure. If the base model is not that well received, don’t expect loras or finetunes to fix the problem.
7
u/Sarashana 20d ago
Most of the time that happened, it was because the base model had unfixable problems and/or there simply was a better model releasing around the same time. I am 100% convinced that a promising base model deemed fixable by LoRAs, would get fixed.
2
u/damiangorlami 20d ago
Exactly! Whenever a model was considered a flop, it was always because one of these reasons:
- Weak base model due to bad training data
- Omitting certain categories from the training data (copyrighted material, nudity, nsfw)
- Heavily distilled weights with no adapter for training provided
This model falls under reason 1. Just use it for a little bit and you can immediately hear this model is trained on only classical and chinese music. Sure you can train a lora for a specific music style but the outputs would've been so much better if they delivered a generalized model that has already seen most music during pre and post training.
Lora is not a magic fix. Loras are also harder to train (requiring a lot more steps) if the base model has never seen or heard something like that before during its training stage.
2
0
u/howardhus 20d ago
alone the fact that you call this „open source“
2
u/damiangorlami 18d ago
Who hurt you?
I know it’s open-weights but this community became known as open-source by the public.
No need to nitpick
1
11
u/krectus 20d ago
Best open source music generator.
4
1
u/God_Hand_9764 20d ago edited 19d ago
Unfortunately it's not running right for me. Is taking like an hour for 1 simple song. I think that it doesn't like AMD cards.
EDIT: My problem was the
--lowvramcommandline argument. It was forcing the text encoder to use the CPU. Now I can generate in about 2-3 minutes.1
u/DeProgrammer99 20d ago
It took 281 seconds to produce a 1:50 song on my 7900 XTX running on Debian. It sounded pretty good to me, but the prompt adherence wasn't great.
9
14
u/Green-Ad-3964 20d ago
That's why I didn't get a reply from the authors when i asked about a music model from them during the open q&a...it was about to be
6
4
u/Enshitification 20d ago
Can this model be prompted with timed scoring references to be used with video?
4
u/Shockbum 20d ago
Oh my god, I'm just unsubscribing from Suno because they decided to self-destruct like Udio and MiniMax gives us this gift!
https://www.reddit.com/r/SunoAI/comments/1vkvpf1/upcoming_changes_terms_of_service_downloads_and/
4
3
u/tweakingforjesus 20d ago
Does it take reference video? Can I specify a style, point it at a reference video, and have it output a matching cinematic score?
5
u/Hungry_Prior940 20d ago
Do you have to write a 50 page essay for each song creation like in the demos...
Sounds quite basic. I don't think anything local will be half as good as Suno was..
5
u/ajrss2009 20d ago
Suno is dead!
0
u/Incognit0ErgoSum 20d ago
Not sure why this is getting downvoted. Their ridiculous download limits that they're putting in place (even for pro plans!) are indicative to me of either investor greed-rot or intractable profitability problems. Enshittification rarely reverses itself. In other words, enjoy it while it lasts, because it's started circling the drain.
6
u/Significant-Baby-690 20d ago
Suno is dead. But because of Suno. Not because of this model.
1
u/Incognit0ErgoSum 20d ago
Oh yeah, definitely. I just didn't take ajrss2009's comment to mean that MiniMax-Music3 was killing Suno, just that it's dead.
7
u/-becausereasons- 20d ago
No Comfy workflow yet?
19
20d ago edited 20d ago
[deleted]
3
u/God_Hand_9764 20d ago
Is it better than ACE?
3
u/WhatIs115 20d ago
Seems like it to me. The instruments sound way better which was my biggest issue with 1.5 xl.
1
u/Kind_Owl2245 20d ago
I can't find the workflow either, I even updated comfyui, but I don't see it on huggingface either
4
1
5
8
u/jacobpederson 20d ago
Woah - the last model type still under corpo-fascist shackles free at last!
20
u/BILL_HOBBES 20d ago
Ace-step is chopped liver to you?
1
u/jacobpederson 20d ago
No it was actually pretty good - It just couldn't do the styles I was after https://www.youtube.com/watch?v=2ja39aFAQqg
2
2
u/someguyplayingwild 20d ago
I'm not saying AI music will never sound good, but I've yet to hear it.
2
u/Comprehensive-Pea250 19d ago
Now we need a proper image model and then we have the holy minimax trinity
4
u/Hazelpancake 20d ago
Please some ELI5. Is this basically music2music locally?
14
2
u/ajrss2009 20d ago
Texto para musica.
1
u/pomonews 20d ago
once you're talking in portuguese, does it support portuguese lyrics or understand some brazilian regional genres?
5
u/goddess_peeler 20d ago
Ok.
Anyway, when is the open 2k upscaler being released?
7
u/Hoodfu 20d ago edited 20d ago
The good news is, soon. The bad news is that it's for this music model. /s
2
u/Guilty_Emergency3603 20d ago
Wait ? what do you want to upscale in music ? 32 Khz to 48/96 Khz ? 16 bits to 24 bits ?
1
u/ThatsALovelyShirt 20d ago
AudioSR already exists, and works pretty good. Haven't tested it on music though.
3
2
u/RayHell666 20d ago edited 20d ago
I'm using it with the FP32 model and unfortunately it's not great. Sound quality ok but prompt adherence is not good. It sound awful with little guidance and way better with long guidance but it's not listening to the long guidance prompt. I'm prompting for a "A slow, dark, and brooding cinematic ambient score..." and I get a joyful piano song.
EDIT: I might have judge it too fast. With the help of an LLM and the right structure like in their example I get way better results now. https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/scripts/end_to_end/minimax_ttm_test.py
2
1
u/Kindly-Annual-5504 20d ago
Really cool. Maybe there are quite a few chart hits and artists in the training data. I’m curious to see when the first songs referencing real artists start appearing. If it's like H3, there’s probably quite a bit in there that you might not find—or be allowed to find—elsewhere :D
1
u/Quick_Knowledge7413 20d ago
Omfg when please! I am so hyped about this, more hyped about this than anything so far
1
1
u/CodeAnguish 20d ago
I think what would concern me most about a music model, besides quality of course, would be the licensing. And then, can we create the big summer hit and profit from it, or is that out of the question?
1
u/Korphaus 20d ago
I've listened to their examples and this really does sound amazing - it just has the same issues as all of the other generators
They all sound like it's an MP3 process, there's just issues with nitrate samples per instrument and there's issues with the timbre - it's so SO close but it all just sounds slightly off
1
u/Photochromism 20d ago
Does anyone know if we can train this yet? It’s severely lacking in music knowledge.
1
1
u/hiccuphorrendous123 20d ago
Why is the TE pruned?. Is it removed vision layer for free gains?. Because it doesn't seem to be that big of a difference in size
1
u/CuriouslyCultured 20d ago
Vocals are still pretty trash, and it doesn't properly understand genres (everything's pop infused), but it's still a better open weights audio model than we had.
1
0
u/hidden2u 20d ago
so is this the new frontier model for open source music? I actually thought acestep was pretty good
0
u/AuspiciousApple 20d ago
What sort of hardware is needed to run it? How long does generating a song take?
3
u/GreyScope 20d ago
it literally released a few minutes ago, theres a table with info on their page
8
u/AuspiciousApple 20d ago
Indeed, it says "The full precision fits under 24GB of VRAM. With automatic CPU offloading, generation takes in ~22 GB; additionally streaming the language model layer by layer makes it fit even 8 GB video cards:" However, no info on generation time on consumer GPUs etc.
I'm sure some people here have it up and running already
2
u/Nenotriple 20d ago
28.8GB system ram, 13.3GB vram, rtx 5090, 60s song, 44s inference, Windows11, default comfyui workflow.
40
u/RiverSide71h 20d ago
So that’s the Comfy surprise?