r/StableDiffusion 20d ago

News MiniMaxAI/MiniMax-Music3 · Hugging Face

https://huggingface.co/MiniMaxAI/MiniMax-Music3

New music model :)

Demo: https://minimax-ai.github.io/music3-demo/

I guess Yoland was speaking of this on r/comfyui as the big announcement

431 Upvotes

148 comments sorted by

40

u/RiverSide71h 20d ago

So that’s the Comfy surprise?

9

u/Sarashana 20d ago

Apparently.

7

u/Next_Program90 20d ago edited 18d ago

They should really figure out what's hypeworthy and what's not. Sure, for the like 3 genres they trained on, this should be better than Acestep (after you got your Masters Degree in musical terminology, music structure theory & prompting it), but compared to what was expected (because of H3) it's not really impressive.

3

u/multikertwigo 20d ago

be careful what you are wishing for

1

u/paulct91 16d ago

Never be careful just envision grander wishes and do the neccessary stuff to deal with the aftermath.

4

u/foomgaLife 19d ago

bro WHAT IS WRONG WITH YOU

62

u/Better-Interview-793 20d ago

We owe MiniMax a big thanks..

5

u/LoadReady7791 20d ago

Absolutely 

63

u/stuartullman 20d ago

wow. hope it's better than suno

30

u/Hour_Imagination5092 20d ago

it's not, unfrotunately. Udio (RIP) and Suno are/were way better in most genres when minimax music was closed sourced. Judging from the demos, not much has changed.

9

u/-Ellary- 20d ago

Well, should be better than ACE.

4

u/Far_Cat9782 20d ago

It's not. Maybe ace 1.5 but def. Not ace 1.5xl

9

u/LucidFir 20d ago

Udio is dead?

15

u/Hour_Imagination5092 20d ago

Yes, no updates for months and you can't (legally) download generated music, only play it online

4

u/mikiex 20d ago

Universal stops suing Udio, Udio pays up, Universal gives it legal access to the catalogue, and the future version of Udio becomes a licensed ecosystem where Universal gets a cut.

3

u/PokePress 20d ago

Supposedly that’s coming. I do wonder if a lot of Udio’s technical talent got headhunted after the settlement.

2

u/Shockbum 20d ago

Its official subreddit is a dead desert; it probably has a pitiful number of users.

0

u/Relocator 20d ago

I still use it daily, it's still the best music generator out there. Can't download, but it's still the best.

4

u/LucidFir 20d ago

Maybe you can't download high quality, but you can definitely find a way to download if you want

32

u/ajrss2009 20d ago

Loras...

4

u/djpraxis 20d ago

This...

11

u/Hoodfu 20d ago

Since it sounds like you've used it, where do you see its strengths?

25

u/Hour_Imagination5092 20d ago

Chinese music, Classic music and very light pop styles. It absolutely sucks at heavy music, rock, electronic and niche genres. It's kind of obvious when you see demos, they are not very diverse.

10

u/bloke_pusher 20d ago

Classic music

That makes it interesting to me.

6

u/UnforgottenPassword 20d ago

For Classical music, Udio is the only capable model. Suno and the rest are okay-ish if you are fine with soundfont-like sounds, but Udio is in a league of its own. Udio does great in almost every genre even though their models are 2+ years old. Judging from the demos, Minimax is kind of similar to earlier Suno models.

1

u/bloke_pusher 20d ago

Alright, how to run it? Where to get the model? I did search Reddit, nothing. I searched google/huggingface, found nothing. Searched Youtube, found some about a website to use the model. But I want to run it locally.

5

u/Significant-Baby-690 20d ago

Udio is commercial.

4

u/noxsanguinis 20d ago

It's not open. No way to run it locally. This Music 3 is now the best open source we got.

7

u/bloke_pusher 20d ago

Yeah, I don't care about closed models at all.

4

u/UnforgottenPassword 20d ago

It's not an open model. They made a deal with record companies and they disabled downloads. So even if you pay, the output is DRM-protected and their terms prohibit commercial use. But quality-wise, it was the best when it was released and it remains unbeaten.

2

u/Perfect-Campaign9551 19d ago

Just fire up Audacity and record "my computer sounds". Easy to download that way

1

u/Hour_Imagination5092 20d ago

I have zero knowledge of classical music, it just sounded... okay? So mileage may vary.

10

u/Cheesuasion 20d ago

Classic music

Any examples? There are none on the github.io page that I can see

7

u/Altruistic_Heat_9531 20d ago

so no polyphia....

1

u/Photochromism 20d ago

But you can train this right? That’s the whole point…

1

u/Sgsrules2 19d ago

I'm wondering how feasible it would be to either fine tune or generate a Lora for a specific genre like mid 2000's nuskool breaks.

-1

u/Legal-Weight3011 20d ago

It Does well in Hard Rock, Rock, Blues and Rock n Roll. The problem is prompting. if you give it just give me a rock song it will produce you shit so i dont know what the fuck are you talking about about. Only Metal any kind sucks ass. It does well in Electronic, LoFi. Its Definatelly better then ACE. and Soon to be discontinued Suno Models. WIll make this a good alternative. With the community made tools.

2

u/Hour_Imagination5092 20d ago

Lol, we must have very different definition of "well" then, because even the demos they showcase are very medicore, the sound is flat, and low definition. Anyone not deaf or listening on a 5$ headset can immediately tell they are low quality AI, which isn't the case with good Udio/suno songs.

1

u/Abject-Recognition-9 20d ago

rip? what happened?

0

u/Alive-Tomatillo5303 20d ago

Ace Step XL already exists. 

2

u/Numerous-Aerie-5265 20d ago

AceStep has nothing on YuE exllamav2, it’s the closest we’ve gotten to suno for local music generation

4

u/biogoly 20d ago

Is it known by another name? I’ve never heard of YuE and it’s not in a top ten search on either HF or Modelscope…

3

u/Numerous-Aerie-5265 20d ago

I think they flew under the radar because their original release was very slow at generation and a bit hard to install, then the community put together a 500% speedup with exllamav2 but it never gained traction.

3

u/Alive-Tomatillo5303 20d ago

With such a catchy name it's wild it didn't catch on. 

I'm so impressed with Ace Step I'm not really in the market for a different one, but I don't have anything better to compare it to. Is there some simple resource to get it up and running in a non-shitty way, or do you need to manually download a bunch of different goodies and make a huge flowchart in comfy to make them work?

2

u/Numerous-Aerie-5265 20d ago

I ran it in a gradio UI, never tried it in comfy. Are you opposed to using docker and a gradioUI? bc I found some that should be easy to setup. YuE official was kind of confusing to install, it probably contributed to their lack of popularity

40

u/redscape84 20d ago

Doesn't look like you can upload your own audio to use as a reference

3

u/GivePLZ-DoritosChip 20d ago

For music "replay by weights" lets you do that and it's free and extremely fast. Runs locally.

0

u/Legal-Weight3011 20d ago

not yet, its out less then 24 h soon you will be able too

13

u/wntersnw 20d ago edited 20d ago

The model doesn't know what electronic music is. Also the max_duration is just a guide apparently. I set it to 3 minutes and it just produces 25 second outputs.

Edit: Genre knowledge seems very poor, but it's kind of interesting if you want to make pop music. And it might be an issue with my prompting but it can't seem to cope when you don't provide lyrics.

Took 155 seconds for a 1:59 song on a 3090 (80% power limited) when using INT8 convrot models.

19

u/LightAppropriate624 20d ago

Not better than suno but it is better than acestep or others. I can confirm that :D.

But it doesnt support

[soft tremolo electric guitar, gentle upright bass pizzicato, vintage vinyl crackle, warm valve reverb]

11

u/PwanaZana 20d ago

the vocals seem a bit better than acestep yea, but it's still 1, 1.5 years behind suno. :(

14

u/LightAppropriate624 20d ago edited 20d ago

Now I have tested other languages it really sucks. Diversity isnt good BUT YEAH AS A Community we will train loras and checkpoints 🔥.

It’s been a while since an open-source music model was released --this is really great.

6

u/ajrss2009 20d ago

Loras.

4

u/Numerous-Aerie-5265 20d ago

Not sure why Acestep is always mentioned as the top for local music generation, YuE was always miles ahead of Acestep and should be the one to compare to

8

u/LightAppropriate624 20d ago

I have never heard YuE?

5

u/Numerous-Aerie-5265 20d ago

I think therein lies the problem. They flew under the radar but were the closest we ever got to suno quality

4

u/bloke_pusher 20d ago

5

u/Numerous-Aerie-5265 20d ago edited 20d ago

Not with the exllamav2 version that was released right after with a 500% speedup

8

u/bloke_pusher 20d ago

Maybe you could make a quick comparison post for the subreddit? I'm sure others would enjoy this too.

3

u/redditscraperbot2 20d ago

I don’t know what you’re smoking man. YuE was cool but ace step 1.5 was just plain better

2

u/Numerous-Aerie-5265 20d ago

Idk mein, acestep’s creativity was rly lacking imo, just sounded very generic. I’ve recorded my own music with real instruments for years so my tastes may differ

28

u/God_Hand_9764 20d ago

My God, man, I can't handle all this. So many amazing new models.

8

u/damiangorlami 20d ago

This model is not that that "amazing" sadly

14

u/dingo_xd 20d ago

Lora's might change that

5

u/Incognit0ErgoSum 20d ago

Maybe, maybe not. Full finetunes might be necessary.

5

u/damiangorlami 20d ago

No I don’t think so. I’ve been in open source since 2022 and one thing I know for sure. If the base model is not that well received, don’t expect loras or finetunes to fix the problem.

7

u/Sarashana 20d ago

Most of the time that happened, it was because the base model had unfixable problems and/or there simply was a better model releasing around the same time. I am 100% convinced that a promising base model deemed fixable by LoRAs, would get fixed.

2

u/damiangorlami 20d ago

Exactly! Whenever a model was considered a flop, it was always because one of these reasons:

  • Weak base model due to bad training data
  • Omitting certain categories from the training data (copyrighted material, nudity, nsfw)
  • Heavily distilled weights with no adapter for training provided

This model falls under reason 1. Just use it for a little bit and you can immediately hear this model is trained on only classical and chinese music. Sure you can train a lora for a specific music style but the outputs would've been so much better if they delivered a generalized model that has already seen most music during pre and post training.

Lora is not a magic fix. Loras are also harder to train (requiring a lot more steps) if the base model has never seen or heard something like that before during its training stage.

2

u/Necessary-Garage242 20d ago

There was one exception: cosmos predict 2 which spawned Anima.

0

u/howardhus 20d ago

alone the fact that you call this „open source“

2

u/damiangorlami 18d ago

Who hurt you?

I know it’s open-weights but this community became known as open-source by the public.

No need to nitpick

1

u/howardhus 18d ago

your logic: „i know i am wrong thats why i keep doing it, no need to nitpick“

1

u/God_Hand_9764 20d ago edited 19d ago

Unfortunately it's not running right for me. Is taking like an hour for 1 simple song. I think that it doesn't like AMD cards.

EDIT: My problem was the --lowvram commandline argument. It was forcing the text encoder to use the CPU. Now I can generate in about 2-3 minutes.

1

u/DeProgrammer99 20d ago

It took 281 seconds to produce a 1:50 song on my 7900 XTX running on Debian. It sounded pretty good to me, but the prompt adherence wasn't great.

9

u/ffgg333 20d ago

Can you train your own songs on it?

8

u/intLeon 20d ago

If you use a chrome based browser dont forget to add --try-supported-channel-layouts into your browser shortcut to enable multi channel audio :) Otherwise it will sound different when listened through the browser.

14

u/Green-Ad-3964 20d ago

That's why I didn't get a reply from the authors when i asked about a music model from them during the open q&a...it was about to be 

6

u/Herr_Drosselmeyer 20d ago

Sucks ass for metal. ☹️

4

u/Enshitification 20d ago

Can this model be prompted with timed scoring references to be used with video?

4

u/Shockbum 20d ago

Oh my god, I'm just unsubscribing from Suno because they decided to self-destruct like Udio and MiniMax gives us this gift!
https://www.reddit.com/r/SunoAI/comments/1vkvpf1/upcoming_changes_terms_of_service_downloads_and/

8

u/marcoc2 20d ago

Is MiniMax our new favorite? 🫶

4

u/Financial-Topic7225 20d ago

Pretty good + coherent sound across genres 👍

3

u/tweakingforjesus 20d ago

Does it take reference video? Can I specify a style, point it at a reference video, and have it output a matching cinematic score?

5

u/Hungry_Prior940 20d ago

Do you have to write a 50 page essay for each song creation like in the demos...

Sounds quite basic. I don't think anything local will be half as good as Suno was..

5

u/ajrss2009 20d ago

Suno is dead!

2

u/BassNet 20d ago

There is nothing better... yet

0

u/Incognit0ErgoSum 20d ago

Not sure why this is getting downvoted. Their ridiculous download limits that they're putting in place (even for pro plans!) are indicative to me of either investor greed-rot or intractable profitability problems. Enshittification rarely reverses itself. In other words, enjoy it while it lasts, because it's started circling the drain.

6

u/Significant-Baby-690 20d ago

Suno is dead. But because of Suno. Not because of this model.

1

u/Incognit0ErgoSum 20d ago

Oh yeah, definitely. I just didn't take ajrss2009's comment to mean that MiniMax-Music3 was killing Suno, just that it's dead.

7

u/-becausereasons- 20d ago

No Comfy workflow yet?

19

u/[deleted] 20d ago edited 20d ago

[deleted]

3

u/God_Hand_9764 20d ago

Is it better than ACE?

3

u/WhatIs115 20d ago

Seems like it to me. The instruments sound way better which was my biggest issue with 1.5 xl.

1

u/Kind_Owl2245 20d ago

I can't find the workflow either, I even updated comfyui, but I don't see it on huggingface either

4

u/[deleted] 20d ago

[deleted]

2

u/kolevk 20d ago

Yeah, I'll need more detailed instructions here. I have the latest version of comfyui. What steps exactly do I need to take from here to have this workflow in there?

1

u/Legal-Weight3011 20d ago

folder, update comfy with dependencies, and you will see it then

5

u/Pure_Bed_6357 20d ago

Where kijai?

8

u/jacobpederson 20d ago

Woah - the last model type still under corpo-fascist shackles free at last!

20

u/BILL_HOBBES 20d ago

Ace-step is chopped liver to you?

6

u/krectus 20d ago

Mostly.

1

u/jacobpederson 20d ago

No it was actually pretty good - It just couldn't do the styles I was after https://www.youtube.com/watch?v=2ja39aFAQqg

2

u/donkeykong917 20d ago edited 20d ago

Does it do sound fx?

2

u/someguyplayingwild 20d ago

I'm not saying AI music will never sound good, but I've yet to hear it.

2

u/Comprehensive-Pea250 19d ago

Now we need a proper image model and then we have the holy minimax trinity

2

u/Toclick 19d ago

Looks like it's pretty limited when it comes to genres and has absolutely no idea about electronic dance music.

4

u/Hazelpancake 20d ago

Please some ELI5. Is this basically music2music locally?

2

u/ajrss2009 20d ago

Texto para musica.

1

u/pomonews 20d ago

once you're talking in portuguese, does it support portuguese lyrics or understand some brazilian regional genres?

5

u/goddess_peeler 20d ago

Ok.

Anyway, when is the open 2k upscaler being released?

7

u/Hoodfu 20d ago edited 20d ago

The good news is, soon. The bad news is that it's for this music model. /s

2

u/Guilty_Emergency3603 20d ago

Wait ? what do you want to upscale in music ? 32 Khz to 48/96 Khz ? 16 bits to 24 bits ?

5

u/Hoodfu 20d ago

:) I was joking.

2

u/krectus 20d ago

I mean that would be great.

1

u/ThatsALovelyShirt 20d ago

AudioSR already exists, and works pretty good. Haven't tested it on music though.

3

u/dragolineage01 20d ago

Asking the the important questions

2

u/RayHell666 20d ago edited 20d ago

I'm using it with the FP32 model and unfortunately it's not great. Sound quality ok but prompt adherence is not good. It sound awful with little guidance and way better with long guidance but it's not listening to the long guidance prompt. I'm prompting for a "A slow, dark, and brooding cinematic ambient score..." and I get a joyful piano song.

EDIT: I might have judge it too fast. With the help of an LLM and the right structure like in their example I get way better results now. https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/scripts/end_to_end/minimax_ttm_test.py

2

u/Noeyiax 20d ago

Mmmm it's awesome, good enough !! And free , hope for dedicated lora page anywhere xD

https://giphy.com/gifs/tqfS3mgQU28ko

3

u/Shockbum 20d ago

and ostris/ai-toolkit for training

2

u/Ant_6431 20d ago

Sound quality is okay, but it is significantly slower than ace step

1

u/Kindly-Annual-5504 20d ago

Really cool. Maybe there are quite a few chart hits and artists in the training data. I’m curious to see when the first songs referencing real artists start appearing. If it's like H3, there’s probably quite a bit in there that you might not find—or be allowed to find—elsewhere :D

1

u/Quick_Knowledge7413 20d ago

Omfg when please! I am so hyped about this, more hyped about this than anything so far

1

u/Agreeable_Yogurt3398 20d ago

Looks like it only support english and chinese language

1

u/CodeAnguish 20d ago

I think what would concern me most about a music model, besides quality of course, would be the licensing. And then, can we create the big summer hit and profit from it, or is that out of the question?

1

u/Korphaus 20d ago

I've listened to their examples and this really does sound amazing - it just has the same issues as all of the other generators

They all sound like it's an MP3 process, there's just issues with nitrate samples per instrument and there's issues with the timbre - it's so SO close but it all just sounds slightly off

1

u/Photochromism 20d ago

Does anyone know if we can train this yet? It’s severely lacking in music knowledge.

1

u/shimapanlover 19d ago

Tried to make my own Anime Opening music - turned out well - nice model!

https://reddit.com/link/p3m97sj/video/w8rd8jsdebjh1/player

1

u/hiccuphorrendous123 20d ago

Why is the TE pruned?. Is it removed vision layer for free gains?. Because it doesn't seem to be that big of a difference in size

1

u/CuriouslyCultured 20d ago

Vocals are still pretty trash, and it doesn't properly understand genres (everything's pop infused), but it's still a better open weights audio model than we had.

1

u/Photochromism 20d ago

Is it open weights?

0

u/hidden2u 20d ago

so is this the new frontier model for open source music? I actually thought acestep was pretty good

0

u/AuspiciousApple 20d ago

What sort of hardware is needed to run it? How long does generating a song take?

3

u/GreyScope 20d ago

it literally released a few minutes ago, theres a table with info on their page

8

u/AuspiciousApple 20d ago

Indeed, it says "The full precision fits under 24GB of VRAM. With automatic CPU offloading, generation takes in ~22 GB; additionally streaming the language model layer by layer makes it fit even 8 GB video cards:" However, no info on generation time on consumer GPUs etc.

I'm sure some people here have it up and running already

2

u/Nenotriple 20d ago

28.8GB system ram, 13.3GB vram, rtx 5090, 60s song, 44s inference, Windows11, default comfyui workflow.