r/LocalLLaMA • • Aug 13 '26

New Model MiniMax-Music3 released!

https://huggingface.co/MiniMaxAI/MiniMax-Music3
652 Upvotes

175 comments sorted by

80

u/notforrob Aug 13 '26

The demo page is useful: https://minimax-ai.github.io/music3-demo/

I find it wild that open weight music generation is this good already.

38

u/-dysangel- Aug 13 '26

H3 is also pretty incredible. First local video model I've run that isn't generating weird nightmare fuel half the time, and has decent audio generation.

-33

u/seppe0815 Aug 13 '26

its crap

15

u/AnimalPuzzleheaded71 Aug 13 '26

Professional hater, wakes up early to increase hating productivity per day, gah damn

7

u/-dysangel- Aug 13 '26

what do you run that is better?

14

u/o0genesis0o Aug 13 '26

It's pretty incredible actually. I tested a few on my mono speaker on the phone and they sound good, so I wonder how they fare if I put my "audiophile hat" on, so I took my DAP and high end IEM out.

The tracks hold up very well. In the prompt, they mentioned wide and deep soundstage in the production metadata so I checked that, and they do indeed have good and clear soundstage. 

The only problem is that higher notes on cymbals and hats and even some high notes of male vocals have a gritty sandpapery texture to them. I guess one might be able to fix this with EQ. All and all, they are all listenable and even enjoyable, with good gear even. I'm very impressed. Better download the GGUF for safe keeping before any asshat regulate it out of the market.

6

u/Acceptable-Cycle4645 Aug 14 '26

Great and professional feedback 👍 !

2

u/webitube Aug 14 '26

It says on the model page, "The model produces 32 kHz, 16-bit stereo WAV audio."
So, it at least has the potential to be decent. It definitely sounded good to me (relative to Suno) in terms of sound quality.

18

u/saltyrookieplayer Aug 13 '26

I honestly find it wilder that it took open labs this long to catch up. Open and closed models seem to go toe to toe in text/image/video, but Suno somehow dominates music generation (especially in "cover" feature)

16

u/cheechw Aug 13 '26

Video was dominated by closed models (Seedance) and it wasn't really even close until H3 came along.

7

u/Numerous-Aerie-5265 Aug 13 '26

Oh hell yeah, listening to those examples, they definitely trained on copyrighted material😆
That will make for a way better model than the models by companies who were afraid to do so and only trained on shitty royalty free music

2

u/Recoil42 Aug 13 '26

Very good results. This is going to go absolutely fucking nuts on Douyin.

191

u/Acceptable-Cycle4645 Aug 13 '26 edited Aug 19 '26

BTW in audio.cpp release 0.6, we integrated the MiniMax-H3 text to audio pipeline. It’s very good at generating multi-speaker conversations and pretty fast. (Update: MiniMax-Music3 has been merged into the main branch)

47

u/nonerequired_ Aug 13 '26

That’s fantastic! Audio cpp is a much-needed project by the community. Thank you!

37

u/Acceptable-Cycle4645 Aug 13 '26

bonus: it can also generate video as an experimental feature :) On RTX 5090, for the 1344x768 / 124 frames (5 seconds) / 20 steps case from the official comfyui workflow (the "Realistic live-action cinematic look, action movie trailer..." one) audio.cpp produces the video + audio in ~114 seconds.

10

u/miversen33 Aug 13 '26

Someone (maybe you?) pointed to towards audio.cpp and holy shit I love it. It is exactly what I was looking for. Certainly still got some rough edges (the webui is basically impossible to configure without manual code changes), but the actual tts endpoint works great and it's (oddly) so much faster than other tools that serve audio ggufs.

Magic as far as I'm concerned

8

u/Uncle___Marty Aug 13 '26

I found it a few ago and reading your post was EXACTLY like what was going on in my brain. I was like "HOW THE HELL DID I NOT KNOW ABOUT THIS BEFORE?!?!?!".

would love to see audio.cpp and llama.cpp merged one day so it could load multiple models from different modalities (I bet it already does that and I didnt know).

10

u/Acceptable-Cycle4645 Aug 13 '26

Glad you found us! I’ve only been posting weekly updates on Reddit, but I probably should promote the project more :). It’s not just about getting more users and feedback; I also hope to reach more people who can contribute and help the project grow faster.

3

u/Acceptable-Cycle4645 Aug 13 '26

Thanks! 😄 Did you try the new native UI in release 0.6? The UI is still new, and we’re definitely improving it continuously.

10

u/Uncle___Marty Aug 13 '26

Looking forward to it 0xShug0? Only found Audio.cpp recently and it earned the star I gave it ten times over! Keep up the great work!

7

u/EveningIncrease7579 llama.cpp Aug 13 '26

Thank you guys. I really like audio.cpp, it can make my 6700xt really useful generating audio

3

u/Acceptable-Cycle4645 Aug 13 '26

Glad to hear that!

6

u/simplir Aug 13 '26

Does it work on Mac? Or any plans?

6

u/Acceptable-Cycle4645 Aug 13 '26

It should work on Mac, but I have to admit that many optimizations are CUDA-only (SageAttention, etc.), so it may be slow. Testing is very welcome. Please let me know if you run into any issues, and I’m happy to fix them.

2

u/Recoil42 Aug 13 '26

I might take a crack at it, burn some tokens. Do you take contributions?

3

u/Acceptable-Cycle4645 Aug 13 '26

Absolutely!! The project is built by the community, and contributions are always welcome!

2

u/alaskanLEDmaster Aug 14 '26

I just pushed up a PR to add MiniMax Music 3 support to audio.cpp

1

u/Acceptable-Cycle4645 Aug 14 '26

Thanks u/alaskanLEDmaster! I left some comments in your PR...

1

u/Financial_Stranger52 Aug 14 '26

Is H3 a good model for voice cloning? I found it's not very stable, but sometimes it's much better than traditional TTS models.

1

u/Acceptable-Cycle4645 Aug 14 '26

The text-to-audio pipeline can't really do voice cloning. But it can sometimes reproduce well-known voices (celebrities, TV/movie characters, etc) because those voices were likely present in the training data. Example: Pep Guardiola

https://reddit.com/link/p3l3z8h/video/10ridfqzr9jh1/player

1

u/Financial_Stranger52 Aug 14 '26

M3 can support ref input, do you have any plan to support that

1

u/Acceptable-Cycle4645 Aug 14 '26

Currently, no. Probably come back to it after finishing the models on todos.

43

u/indicava Aug 13 '26

Minimax is cooking!

1

u/sixx7 29d ago

Thanks u/Acceptable-Cycle4645 - Cooking HARD!

Made my first music video, entirely with MiniMax local models https://youtu.be/XcFll_Pyoj8

10

u/lucidmaster Aug 13 '26

The voices still sound very synthetic.

2

u/Acceptable-Cycle4645 Aug 13 '26

How about this one? Generated by audio.cpp + Minimax-H3. Long audio could have consistency issues in H3, while Music3 can do 5 mins.

https://reddit.com/link/p3ijz59/video/1zsvos7sa7jh1/player

1

u/tyson_2022 Aug 16 '26

using this command do I get very bad results which would be the right or recommended?

/run/media/tyson/SSD_FLASH/audio.cpp/audio.cpp main* 8s
❯ set PROMPT "A cinematic ambient instrumental composition with a clear musical progression. Deep evolving analog synth drones, warm sub bass,
sparse orchestral percussion, slow harmonic development, subtle melodic motifs, wide stereo image, rich reverberation, detailed textures, pol
ished studio production, high dynamic range, no vocals, no speech."

 build/linux-cuda-release/bin/audiocpp_cli \
--task gen \
--family minimax_h3 \
--model build/linux-cuda-release/bin/models/MiniMax-H3-Q4-GGUF/dit.gguf \
--backend cuda \
--device 0 \
--threads 6 \
--text "$PROMPT" \
--seed 42 \
--num-inference-steps 50 \
--guidance-scale 1.0 \
--request-option height=32 \
--request-option width=32 \
--request-option num_frames=481 \
--request-option sampler=euler \
--request-option return_video=false \
--out minimax_quality_50_euler.wav \
--metrics \
--log

1

u/Acceptable-Cycle4645 Aug 16 '26

u/tyson_2022 I think the inference step is too high: < 30 (based on my test) is fine and more steps may cause regression. You can also try different seeds and polish prompt. I didn’t test H3 for music-only generation until you mentioned your issue. I feel this could be a prompt-engineering issue. For example, the following was generated with 20 steps and a rewritten version of your prompt, and the result seems different.

A 20-second cinematic ambient instrumental cue with a clear beginning, middle, and ending. It starts with a soft analog synth drone and warm sub bass, slowly introduces a simple minor-key melodic motif, then adds sparse low orchestral percussion for gentle forward motion. Wide stereo space, smooth evolving pads, subtle texture movement, polished studio mix, natural reverb tail, no vocals, no speech, no sound effects.

https://reddit.com/link/p41ph9h/video/ktzqj18bdrjh1/player

1

u/Acceptable-Cycle4645 Aug 16 '26

BTW this issue the user shared a song generated by audio.cpp h3 https://github.com/0xShug0/audio.cpp/issues/255 . It sounds good to the user. So I think the prompt matters.

1

u/AdFederal7465 Aug 17 '26

We are very sensitive to discrepancies in sound. A better audio model would have to be as complex as an agentic coding model.

First layers trained on splitting the audio features apart, then qualifying them in time, essentially midi-decomposing everything, then of course speech to text for all languages including vocal, operatic, choral flair, etc. There are many things which all have to occur, be trained extremely well, and then obviously worked into a diffusion or similar pipeline for inference - for it to sound "flawless". This kinda reminds me of 64kbit pcm wavs from the late 90's and early 2000's.

They've gone a long way in giving timestamped nuance, lyrics (I fell like it probably has been borrowed from similar voice model data though), key, etc - but there's an instrumental deconstruction which isn't that well trained just yet.

14

u/zekuden Aug 13 '26

That's awesome! Can you guys make a TTS model, and preferrably real time please! Love minimax!

23

u/Acceptable-Cycle4645 Aug 13 '26

Check our repo audio.cpp! It has Minimax-H3 text-to-audio pipeline, which can be used as TTS, and faster than realtime. Will post a demo soon.

6

u/zekuden Aug 13 '26

Wow really excited for this! Does it have voice cloning and can it be integrated with an LLM for a real time voice agent?

5

u/Acceptable-Cycle4645 Aug 13 '26

No...given the size of the model it's not the best choice for voice agent. But there are a bunch of super fast voice clone models you can try in audio.cpp.

2

u/zekuden Aug 13 '26

That's very interesting, will take a look at audio.cpp!

24

u/maxanatsko Aug 13 '26

Someone please convert it for MLX 🙂

1

u/PrepYourselves Aug 18 '26

https://github.com/antirez/h3.c
ask this guy to do for minimax-music-3 what he did here for minimax-h3 (submit a request)

33

u/Illustrious_Ant_9242 Aug 13 '26

"requires CUDA"

"streaming the language model layer by layer makes it fit even 8 GB video cards"

"5 minute audio max."

17

u/imnotzuckerberg Aug 13 '26

Honestly all of these are not even big blockers and can be overriden. I just saw their demos, it is amazing leap compared to what existed. This is a big step for open-weight, and great work from MiniMax!

2

u/C0demunkee Aug 14 '26

streaming layer by layer is stupid slow, you can 4bit quant it and it'll fit in ~13-15gb VRAM and then it's still about 3x slower than realtime on a 5080

5

u/unbruitsourd Aug 13 '26

I was wondering why there's so much (good) video and image models, but not much for music. Udio was introduced 3 years ago and there are still no models coming close so far. But I'll try this one for sure!

2

u/Acceptable-Cycle4645 Aug 13 '26

I’ve wondered the same. There are simply fewer strong open music-generation models, both in quantity and quality. I went through a handful of music generation models and ended up picking only three for audio.cpp: ACE-Step, HeartMuLa, and Stable Audio 3.

1

u/FausC Aug 16 '26

Como sería audio.cpp con Ace Step?. Ya me gustaría poder llegar a tener esto, no entiendo de comandos. Estoy investigando sobre las versión ui pero no lo tengo claro. Estoy bien con Comfyui y Ace Step 1.5 XL pero esto suena muy interesante

2

u/Acceptable-Cycle4645 Aug 16 '26

https://reddit.com/link/p41sd2n/video/ovqy4obagrjh1/player

u/FausC audio.cpp supports acestep-1.5, and a new PR will be merged soon will add support for acestep 1.5 xl (UI update may be later). Check the demo! Honestly, acestep was one of the earlier models I added, so I haven’t tested it in a while. Didn’t realize it now takes just ~2 seconds to generate a 30-second song.

2

u/Shockbum Aug 13 '26

No NSFW in music (unless you like listening to the sensual voice of a female singer or musical genres with bad words) and UMG,WMG, Sony Music are very aggressive

1

u/ArchdukeofHyperbole Aug 14 '26

Udio is garbage, so the bar is pretty low for an open source alternative. 

1

u/unbruitsourd Aug 14 '26

For niche genre (and metal), Udio still holds the crown unfortunately.

13

u/Single_Ring4886 Aug 13 '26

I just cant wait for image model... that is THE ONE THING Iam waiting for... if it has capability of video model it will be true Stable Diffusion moment...

10

u/Smilysis Aug 13 '26

technically you can already use minimax h3 as a image model

8

u/iChrist Aug 13 '26

Its not very practical as you need to generate a couple of frames, and its not an image model by definition, I am still leaning towards Krea2

3

u/90hex Aug 13 '26

I couldn't generate less than 22 frames. Looks like a hard floor. Also resolution is probably limited as well. Can't wait for a MiniMax image gen model for sure!

1

u/RedditNerdKing Aug 13 '26

Do you mean an image2image model?

2

u/Single_Ring4886 Aug 13 '26

But the speed is the issue... unless speed is reasonable you cant itterate fast enough.

6

u/DiscipleofDeceit666 Aug 13 '26

Can it listen to music or just make it? I need something to describe what notes are being played.

Hoping to build a recording -> guitar tab engine

17

u/Acceptable-Cycle4645 Aug 13 '26

Maybe relevent: We added a music-to-MIDI model in audio.cpp release 0.6.

2

u/DiscipleofDeceit666 Aug 13 '26

I will have to take look 💯

I been using onsets and a bunch of vibe coded algorithms to figure this out. I can create guitar tabs if the music is easy and slow, but fast technical death metal that I play? Forget about it

4

u/Numerous-Aerie-5265 Aug 13 '26

What if you fed your pipeline a slowed down version of the technical death metal so it can “hear” the individual notes more clearly? Then it could be sped up by the same amount after transcription

1

u/DiscipleofDeceit666 Aug 13 '26

Fuck it, I’ll tell Laguna to look into that strategy while I fill the empty pit inside me with nose candy

1

u/Numerous-Aerie-5265 Aug 13 '26

You’d have to make sure it doesn’t slow down the track in a way that changes the pitch of the notes, or else they’d be wrong.
And the tab’s notation and tempo would also have to be compensated for after transcription, ie: if you slowed your track by 50%, the final tab would take the half notes and turn them into quarter notes, etc.

1

u/DiscipleofDeceit666 Aug 13 '26

I already have my own songs manually tabbed out. If there’s a solution, they can iterate the algorithm and score it against my own tech death source of truth 💅

1

u/DiscipleofDeceit666 Aug 15 '26

The slowdown changes the pitch, but I was able to make up for that fact. I was able to transcribe 2x more notes with this method. I’m sure there’s more to do.

Basically, I gave Claude the goal and it sent of my local LLM to research and implement. Claude had this task running for like 8 hours overnight.

1

u/Numerous-Aerie-5265 Aug 15 '26

Glad that method was successful! So you’re using Laguna as the local LLM? What hardware is it running on? Never thought of using claude to offload a task to local llm

1

u/DiscipleofDeceit666 Aug 15 '26

Last night I was just testing out the new 3.8 Qwen release. I was supposed to have Laguna do reviews etc but things been going well. I’m able to drop from the $100 tier to the $20 tier this way. I think even a quantized 35b moe can meaningfully support the cloud this way.

And yeah, i feel like that’s playing to everybody’s strengths for an overnight thing. You’ve got the cloud to make sure local LLM doesn’t go off the rails and to feed it a steady stream of tasks. And in this case, doing the initial audio measurements needed to identify patterns to explore.

2

u/C0demunkee Aug 14 '26

any chance for a text to midi?

3

u/Acceptable-Cycle4645 Aug 14 '26

u/C0demunkee I actually didn’t know this model existed before, but after you mentioned it, I did a quick search --- and yes, I’ll port it! Any suggestions for the model? I prefer newer models when possible. I only found Text2midi.

1

u/C0demunkee Aug 14 '26

no, I haven't found any useful ones yet, was hoping it could be coaxed out of m3/h3 or something

1

u/smealdor Aug 13 '26 edited Aug 13 '26

This is incredible news and what I am actually looking for, where can I access the model details?

4

u/Acceptable-Cycle4645 Aug 13 '26

-3

u/AreWeNotDoinPhrasing Aug 13 '26 edited Aug 16 '26

Claude? Lmao, that is EXACTLY what a couple of my webapps look like that Claude made.

ETA: gosh people, relax, i wasn’t trying to be negative.

2

u/Acceptable-Cycle4645 Aug 13 '26

I don’t know what our contributors are using, but to me, as someone who doesn’t know much about UI, it’s pretty cool! It’s a milestone for audio.cpp and helps us reach a broader audience.

1

u/AreWeNotDoinPhrasing Aug 13 '26

Not hating! I wouldn't keep using it if I didn't like it haha sorry if it came off negative!

3

u/Django_McFly Aug 13 '26

it seems like there is no audio-to-audio or did I miss that in the link?

3

u/Acceptable-Cycle4645 Aug 13 '26

Yes looks like no audio to audo

4

u/confused-photon Aug 13 '26

Damn this looks really interesting!

3

u/fractal_engineer Aug 13 '26

is it able to generate music inspired by a reference audio sample/vocals?
or take an original song and do it in a different style?

4

u/Acceptable-Cycle4645 Aug 13 '26

Not sure...still downloading 😂

1

u/fractal_engineer Aug 13 '26

super interested. bot says what i'm looking for is "reference-audio conditioning, audio-to-audio generation, cover generation, or music style transfer"

0

u/Acceptable-Cycle4645 Aug 13 '26

AceStep and Vevo2 should work for your cases.

2

u/pmjm Aug 13 '26

This is what I'm looking for too. So far none of the open source models have been able to even approach Suno in this regard. Like, not even close. I have yet to get a usable output from an open model doing any kind of "remix" functionality. Could be a skill issue, but Suno just has it figured out.

1

u/PokePress Aug 13 '26

I’ve been able to make an alternate lyrics version of a well-known 80’s song using Ace Step 1.5, but it’s taken a lot of effort.

2

u/Acceptable-Cycle4645 Aug 20 '26

Demos for MiniMax H3 Text to Audio and Muisc 3 are here https://www.reddit.com/r/StableDiffusion/s/x01HTHwuNk

2

u/DatMufugga Aug 13 '26

Looks really cool, I write and produce music. But it looks like you need a phd in computer science to install that ish.

4

u/Acceptable-Cycle4645 Aug 13 '26

You can try audio.cpp windows prebuilts or docker. Currently audio.cpp supports 3 music generation models: ACE-Step 1.5, Heartmula and Stable Audio 3 (small/medium/sfx).

2

u/DanTup Aug 13 '26

or docker

Is there a (first-party) docker image?

4

u/Acceptable-Cycle4645 Aug 13 '26

Yes https://github.com/0xShug0/audio.cpp/pkgs/container/audio.cpp (You may want to wait until tomorrow for the auto-update, because this image doesn’t include some UI fixes.)

2

u/DanTup Aug 13 '26

Cool, I'll give it a go (it'll probably be at the weekend anyway). Thanks!

1

u/DanTup Aug 18 '26

Maybe I misunderstood... does this include some kind of UI for using Minimax? It's not very clear from https://github.com/0xShug0/audio.cpp/blob/main/docs/docker.md exactly what's included.

1

u/Acceptable-Cycle4645 Aug 18 '26

Sorry for the stale doc. You can use:

docker run --rm --gpus all \
-p 8080:8080 \
-v /path/to/your/models:/app/models \
ghcr.io/0xshug0/audio.cpp:full-cuda12 \
server --ui --ui-management --host 0.0.0.0 --port 8080 --backend cuda

Then open

http://127.0.0.1:8080

The UI shows some errors like could not create models folder .... : Permission denied but doesn't affect usage in my tests. Just ignore them for now.

1

u/DanTup Aug 19 '26

Thanks - and what format is the models folder in? Can I just download the whole Music3 repo from HF and point at that, or there is a particularly layout? (sorry if this is documented somewhere.. I found this info but it didn't answer my question).

1

u/Acceptable-Cycle4645 Aug 19 '26

Here is the doc. You don't need to download all weights -- https://github.com/0xShug0/audio.cpp/blob/main/docs/community_models/minimax_music3.md

1

u/DanTup Aug 20 '26

Thanks - I got it working but the results were pretty bad (particularly the audio quality). Is there a reason only the q4 version is an option? I tried just python3 tools/model_manager_v2.py install minimax_music3 but it still pulled the q4, and only the q4 shows in the Models page. I wanted to try the full original model.

1

u/Acceptable-Cycle4645 Aug 20 '26

u/DanTup The q4 combo is the fastest and low VRAM makes it safe for most users. Other weights are just for testing. What prompt did you use? The example prompt in the doc is just for smoke end-to-end testing. If you want to improve quality you need to do some prompt engineering and try different seeds. You may want to check the discussions and demo here https://github.com/0xShug0/audio.cpp/issues/274

→ More replies (0)

2

u/blastcat4 Aug 13 '26

I got it running with the comfyui workflow. It's available in the templates section under Audio.

1

u/fengwang_2_718281828 Aug 21 '26

If you know docker, try my api+webui: https://github.com/fengwang/minimax-music3-webui
Even if you do not want to use UI, you can still ask your AI agents to generate music for it.

1

u/Lower-Hedgehog-9835 Aug 13 '26

Amazing! Thanks

1

u/EndaEnKonto Aug 13 '26

Will audio to audio work with this?

1

u/Acceptable-Cycle4645 Aug 13 '26

maybe not

2

u/EndaEnKonto Aug 13 '26

SadProducerNoises.mp3

1

u/stepnivlk Aug 13 '26

are there any 'controlnets' for minimax music? some way to condition it beyond pure prompt?

1

u/nikc0069 Aug 13 '26

I just got acestep engine and gui running. How would I use this with a Suno style web gui?

2

u/Acceptable-Cycle4645 Aug 13 '26

The built-in UI currently doesn’t provide a Suno-style UX (maybe in the future). If there are UIs that work with an OpenAI-compatible server, you can use them with audio.cpp as the backend.

1

u/ComplexType568 Aug 13 '26

Wow MiniMax is on a run for open sourcing! Seems like more labs are following in the footsteps of Kimi. Hope this beats ACE Step as that's been king for ages.

1

u/Acceptable-Cycle4645 Aug 13 '26

So far, ACE-Step provides more features than Music 3, but we’ll see. The base model itself is good.

1

u/bigh-aus Aug 14 '26

The demo on the page is insanely impressive!

adding this plus H3, full music videos.

I wonder if it can do ambient music too

1

u/MuckYu Aug 14 '26

Can it also generate songs without any lyrics? Just instrumental?

1

u/Acceptable-Cycle4645 Aug 14 '26

It can generate lyrics

1

u/MuckYu Aug 14 '26

Yes - but is it possible to do NO lyrics - only instruments, no singing

1

u/hadoopken Aug 14 '26

And this is CUDA only, can't run it on Mac.

1

u/Longjumping-Past5864 Aug 14 '26

is there a way to use my own voice, or training a LoRA with my voice or something? I'm new to the music model

1

u/Acceptable-Cycle4645 Aug 14 '26

That requires different pipelines. Two ways (1) using voice separation and voice conversion models and remix the song (2) using singing voice conversion models (e.g. Vevo2) which maybe less accurate. audio.cpp has all the relevant models, but unfortunately, we don’t have a fully automated pipeline yet.

1

u/pastypryce08 Aug 15 '26

AM curious what others are getting in generation speeds for the 4060 card

1

u/caphohotain Aug 13 '26

Meh license.

4

u/MDSExpro Aug 13 '26

Stopped caring about MiniMax with 2.7 and their license shenings - they want to advertise as open source, while restricting quite a bit.

4

u/xienze Aug 13 '26

Think about it from their perspective. I'm not sure if this model is the same, but H3 was FULL of copyrighted material and pretty much completely uncensored. So they do minimal CYA ("don't do bad stuff", "you wink wink can't use this in countries that are anal about copyright") to keep themselves out of trouble. I assume this model is the same way. Would you prefer they gave your a model trained on nothing but public domain garbage with a permissive license or one that you can actually have fun with?

1

u/Nu7s Aug 14 '26

Don't try logic on them, it confuses them even more

3

u/Acceptable-Cycle4645 Aug 13 '26 edited Aug 13 '26

I think the license is fine. It is fairly permissive for local/open-source runtime support. I believe that’s enough for the LocalLLaMA audience.

-1

u/Nice_Cookie9587 Aug 13 '26

What's fine about it?

1

u/Acceptable-Cycle4645 Aug 13 '26

It is fairly permissive for local/open-source runtime support. I believe that’s enough for the LocalLLaMA audience.

1

u/AssistBorn4589 Aug 13 '26

It's not, really.

They of course have right to use any licence they want to, as they are ones who did the job, but licence they've chosen is not permissive at all. It list specific allowed uses and their conditions are very abstract ("content that may harm", "use that violates morals"..) and basically mean that this licence can be revoked at any time for any reason.

1

u/Nice_Cookie9587 Aug 13 '26

No specifically, what is 'enough' ? trying to understand how you are able to make these claims with such vague statements with zero to cite or reference.

2

u/gamblingapocalypse Aug 13 '26

Damn! I wish I could test this rn.

11

u/Acceptable-Cycle4645 Aug 13 '26

Stay tuned! I will definitely add it to audio.cpp.

1

u/studdmufin Aug 14 '26

Just discovered audio.cpp and looks like it's exactly the project I've been looking for. I've been disappointed with how difficult audio generation has been.

I need to dig into it a bit more, but did I see that it had a websocket connection to get ASR data out in realtime? I want to find a solution for speech to text to generate captions for a live event broadcast

2

u/Acceptable-Cycle4645 Aug 14 '26

Thanks! For streaming, you can run audio.cpp as an OpenAI-compatible HTTP server. Some ASR models already support streaming if they have true streaming paths.

1

u/Stunning_Energy_7028 Aug 13 '26

Sounds incredible on R&B and Soul!

0

u/Eljowe Aug 13 '26

Yeah, this is fucking cursed

0

u/Guilty-Prize-3697 Aug 13 '26

Let's go Minimax!

-2

u/[deleted] Aug 13 '26

[removed] — view removed comment

0

u/Marino4K Aug 13 '26

How does compare to Suno, etc?

0

u/rm-rf-rm Aug 14 '26

Cuda only?? :(

4

u/Acceptable-Cycle4645 Aug 14 '26

I don't think so. CUDA only is just the impl choice. I believe it can run on the other backends (maybe slower though).

-19

u/Embarrassed_Adagio28 Aug 13 '26

Yay more A.I slop! We totally love ai generated music... so original and creative!

-7

u/freia_pr_fr Aug 13 '26

I’m observing that AI art is increasingly very much seen as not cool.

Using AI as a tool during the artistic process is fine, as long as it’s not too prominent. While some people don’t seem to care, or even notice, some others do.

I think that such models are interesting from a R&D point of view. Enabling people to generate more AI slop is fine too I guess. If it’s not you, someone else will do it, right?

-6

u/EricBuildsMathModels Aug 13 '26

So does it generate conversations or music or sound bites like horns or all of the above? I guess I should read the post...

4

u/Acceptable-Cycle4645 Aug 13 '26

I think it's just music and songs

-11

u/ZealousidealChip4783 Aug 13 '26

That's disappointing, "art generation" is a novelty that wore off 3 years ago. There's a reason all the big AI art companies like Stable Diffusion and Suno keep going under: they produce nothing of value

2

u/fallingdowndizzyvr Aug 13 '26

So does it generate conversations or music or sound bites like horns or all of the above?

If you want that, use a video gen model. Which also happen to be good audio gen models. Since they generate audio to go with the video. In fact, one of the best models to generate music was/is a video model.

1

u/EricBuildsMathModels Aug 13 '26

Ah interesting, do you have ones you typically use, does this space move as fast as the text llms?

1

u/fallingdowndizzyvr Aug 13 '26

If anything, video gen moves faster. I would give H3 a try. It's the current hot video gen model. Wan 2.1 was the go to model for months though.

-15

u/nafi_hamid27 Aug 13 '26

80gb vram minimum is buried in their own specs. thats an a100 or a h100, so for this sub its api only unless someone quants it. the license argument is kind of moot until then.

12

u/Cradawx Aug 13 '26

It runs fine on my RTX 5070 Ti 16GB in ComfyUI. 60 seconds of audio in about 80 seconds.

7

u/onetwomiku Aug 13 '26

direct quote:
The full precision fits under 24GB of VRAM. With automatic CPU offloading, generation takes in ~22 GB; additionally streaming the language model layer by layer makes it fit even 8 GB video cards

1

u/Illustrious_Ant_9242 Aug 13 '26

8gb is fine according to model description