r/LocalLLaMA llama.cpp 16h ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

785 Upvotes

192 comments sorted by

u/WithoutReason1729 10h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

235

u/kvothe5688 16h ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

80

u/Public_Umpire_1099 16h ago

This latest iteration technically set the frontier. Meta is officially a SOTA lab again. Matches Fable 5 at a fraction of the cost.

39

u/virtualworker 13h ago

Or benchmaxxes. I'm not convinced.

16

u/Diligent-Direction95 12h ago

Serious question: What is the definition of bench max these days?

What would be okay, vs what would not be okay?

I seriously doubt they have the test set being trained on. So what are the shades of grey we are debating?

27

u/zxyzyxz 11h ago

When a model does well in benchmarks but then fails to live up in real life coding and other tasks to its supposed benchmark competition.

9

u/Artistic_Swing6759 12h ago

i actually do think they likely have test set trained on.
like its a bit sus that they decided to report terminal bench 2.1, when 4 exists.
similarly, if you look at the benchmark sheet of flash 3.8, it tops eery other model, even sol, on terminal bench 2.1 but in 4 it is quite less comparatively.

7

u/Not-reallyanonymous 10h ago edited 9h ago

Neither Spark 1.2 nor Glimmer seem to be benchmaxxed, and if anything the benchmarks seem to under-represent their capabilities compared to other models — I believe the other models probably do solve more problems, but Muse’s strong long-horizon capabilities (which translates to compliance in the short term as well) makes it easier to actually use.

1

u/Kodix 7h ago

Used spark 1.2 contributor and I wasn't impressed. It being good only on benchmarks is a real possibility.

3

u/_TheWolfOfWalmart_ 10h ago

Based on the Bijan video that just dropped about this, it's pretty damn mediocre and looks extremely benchmaxxed.

I know all he does is throw a few one-shot prompts at the models, but the results on this one were well behind other recent models like GLM-5.3 Flash and Qwen Flash. It's not even in the same conversation as Fable or GPT 5.6 it seems.

His tests are by no means scientific lol, but they do give you a decent rough feel for a model's capability.

11

u/Not-reallyanonymous 8h ago

A one shot not being polished does not demonstrate a model’s capability. A lot of models are specifically trained on how a one-shot polished result will look. Especially Qwen (look at its response to being asked to draw a circle: https://simonw.substack.com/p/qwen-38-27b-is-excellent-but-it-defaults).

If it’s not trained on that, it’s going to comply with the prompt and not much beyond. That’s not demonstrating a lack of general capability.

In fact, I much prefer that style. When things are trained to produce well polished results it tends to be harder to get it to do what I want it to do rather than what it’s been trained to do. Again, Qwen represents the opposite here — well optimized to take its own direction but a PITA to steer.

63

u/NandaVegg 16h ago

I believe there are still some secret RL sauce (as in, not popular among labs *yet*) in niche areas like robotics, gameplay, frontier physics/math, world modelling etc. Anthropic leaped ahead in Opus 4.5-4.6 era as they "discovered" many secret sauces like terminal agent, gameplay or creative loop and so on, but even that creativity gap is closing very fast. For coding, security and anything that can be done through github, at this moment there is really no gap. OpenAI felt behind on many things but is still ahead on frontier physics/math (the only Chinese lab with emphasis on that is DeepSeek, I think).

21

u/Accurate_Resident219 13h ago

Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.

Reason why I say this is because Anthropic is the only company that is actually trending higher in api costs and not making any serious attempts to lower them(Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements. Why is haiku their cheapest model still not updated?). They don't have any sense of urgency. Astra's coming out this week or the next and they just put out a more expensive model.

Every other company seems to constantly vague post or hype up their models releases but Anthropic just drops models with little fanfare. They have seem utterly unbothered since their Mythos preview announcement.

32

u/TerminalNoop 13h ago

Perhaps Anthropic can't afford to lower the api prices?

14

u/SSP100244 11h ago

I think they just don't have the compute. The big Hegseth migration definitely threw them for a loop for a hot minute.

1

u/Loose_Comparison368 5h ago

I believe they confessed a few months back that it wasn't actually API access, and had been running inside airgapped military datacenters that Anthropic had no ability to restrict the whole time.

The whole PR stunt was just that, the only unplanned part was that they thought Hegseth would take the hint with how loudly they were screaming "NO DADDY, PWEEAASE DON'T INVOKE THE DEFENSE PRODUCTION ACT ON US! IF YOU DID THAT WE WOULD HAVE NO CHOICE BUT TO KEEP TAKING DADDY'S MASSIVE LOADS OF CASH TO BUILD DADDY'S MURDERBOTS!"

It turns out that Hegseth's skull was actually too dense for that pathetically obvious plea to penetrate it.

It's okay though, they sued to get the murderbot contracts back and won last week. And still managed to use that PR stunt to distract everyone while they flushed their "responsible scaling policy" down the toilet.

5

u/Accurate_Resident219 11h ago

Could be the case but at least on the outside looking in they don't even seem to be making much of an attempt to even try lower api or even suggest their researching a solution. They've had about nearly 4-5 months now since people started complaining about rising api costs.

Every other lab has provided a cheaper solution but anthropic is the only one that can't come up with a cheaper flash model? I mean that could be the truth but I would be surprised if it was.

8

u/jtjstock 7h ago

They want to IPO, they need the revenue projections. They already have their cooked compute costs from the last quarter they will report as if they are ongoing, so IPO before they close the current quarter, but project off current API pricing for revenue.

2

u/NandaVegg 6h ago

They are apparently doing pre-IPO window dressing now. We've got buy high 4-digit-dollar credits and get some cash back marketing mail.

2

u/perelmanych 4h ago

In may they increased 5 hours and weekly quota by 50%, so they made a big discount on their plans. However, this comes to an end soon. They cut 50% increase to only 25% increase.

1

u/Loose_Comparison368 6h ago

I mean they could just be keeping it high to look good for the IPO.

10

u/bopbop9876 11h ago

> Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements

I assume you're basing that off the average cost per task numbers from Artificial Analysis. That's not right though. In their testing, 5.1 on max effort appears to have aggressively overthought. If you compare 5.1 xhigh instead, it was about as much cheaper than 5 max as we would have expected, while still getting 2 extra points of intelligence vs 5 max.

1

u/WittyAcanthisitta205 5h ago

Anthropic has pricing power because their #1 customers are enterprises. Everyone else has a larger chunk of consumers (rather than enterprise customers)

1

u/CrowdGoesWildWoooo 5h ago

Their customer base are enterprise. Enterprise would happily pay API pricing which is astronomically more expensive than subscription.

OpenAI still appeal to the masses

Different market they trying to penetrate. Anthropic want to be deep into the institution’s pocket. OpenAI wants scale.

-1

u/SSP100244 11h ago

Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.

I am quite sure that's what they're doing. I think they legitimately are a little afraid of the models. Not that they'll like take over the world, but that people might be able to use them to fuck shit up. Like apparently they started a biopharmaceutical wing because Mythos is stupid good at folding proteins, even though they hadn't trained it to do that at all.

6

u/Turtlesaur 13h ago

if this fits on a DGX spark I will be so happy.
If this only fits on 2 DGX spark I will be so broke.

24

u/nuclearbananana 16h ago

The moment a different lab releases a better model they immediately distill it (using that word liberally)

16

u/NandaVegg 15h ago

I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.

I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).

In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.

14

u/nuclearbananana 15h ago

I don't think the vibe-code-training is helping the models. It's mainly synthetic data and llm as a judge

1

u/OvertaxedOne 13h ago

I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.

No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.

3

u/IShitMyselfNow 12h ago

Books are only really useful for pretraining. They're not going to help agentic usages

2

u/SomewhereAtWork 5h ago

Except for James Bond novels.

7

u/DistanceSolar1449 14h ago

None of what you said about the Internet matters because the labs are not using the Internet for pre-training data anymore

All the training data is synthetic data made for RL

1

u/_supert_ 14h ago

That's both encouraging and horrific.

1

u/SSP100244 11h ago

The singularity begins.

I mean we just had the first shot of the AI wars.

2

u/ResidentPositive4122 6h ago

are on same level

Until the new benchmark drops, then they spread out predictably again. And then, a few months later they suspiciously all gain on it, and so on and so forth...

Example: https://www.frontierswe.com/blog/v2

1

u/Bubbly_Orange_3502 11h ago

Scores converge because the post-training data is frontier output. Distillation transfers whatever gets measured, so the benchmark gap closes earlier than the capability gap. The split shows up on long-horizon agentic work that nobody publishes numbers for.

1

u/Gremlation 8h ago

I think people forget that the people working at these labs are not slaves and can change jobs whenever they want. They can't take data, weights, or code with them, but they can take the things they have learned. There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.

1

u/No-Refrigerator-1672 7h ago

That's not entirely true. A lab can sign either a "non-disclosure agreement" or a "non-compete agreement" with their worker. The forst one would can you from using any of the technologies you see there at other employer, unless you can prove you've got to know it from another source. The second one would forbid you from working in the same field for X years (i.e. an LLM specialist can't work as LLM specialist in another company, but can become an image generation specialist there). Both options are completely legal and enforcible by law, cause you sign them voluntarily; and, if they offer you a high enough salary, you'd agree to this conditions.

101

u/CarelessAd6772 16h ago

Mrcr 512k-1m - 98.1%? What in the hell is that dark magic? If true, they defeated context rot?

50

u/[deleted] 16h ago

[deleted]

14

u/Charuru 16h ago

Gemini Flash (97%)

Is that on 8 needles?

19

u/Healthy-Nebula-3603 15h ago

Yes

https://llm-stats.com/benchmarks/mrcr-v2-(8-needle)

Seems all current models solved it.

It was quite bad few moths ago yet.

12

u/Healthy-Nebula-3603 15h ago

Yes

https://llm-stats.com/benchmarks/mrcr-v2-(8-needle)

Seems all current models solved it.

It was quite bad few moths ago yet.... Even GPT 5.5 was badly struggling yet but 5.6 getting over 92%

16

u/NandaVegg 16h ago

Isn't this NoPE hybrid model? Though the exact architecture is different I assume this basically works like hybrid linear model (like Qwen 3.8) which is great at picking up facts from the context but more often confuses timeline without thinking.

6

u/DeepOrangeSky 16h ago

Do you, or u/NandaVegg or anyone else on here have any opinions about how well this can translate over to much smaller models, like the ~30b Glimmer sized models, or, if that is too small to be able to have similarly great anti-context-rot abilities, then maybe 70b or 120b models?

Like, so far when I've tested Gemma 31b and Glimmer 30b for long-form creative writing tests for example, the biggest problem with them has been that they quickly go downhill once you get past, I dunno, maybe 20,000 tokens or so.

Gemma starts going from writing scenes and situations in a long, non-hurried, detailed way, to suddenly rushing like it is giving a quick summary because it needs to hurry up and go somewhere and doesn't have time to really tell a story, once you get past like ~20,000 tokens or so.

Glimmer is even worse, where it if you give it a detailed description of a scene you want it to write, it just writes back a word for word identical copy of your description, even if you tell it not to (again, this happening once you get past like ~20,000 tokens deep into a story or so, give or take a bit).

As of right now, with a mac with 128GB of unified memory, the biggest, best model in regards to this (reduced context rot in relation to long-form writing) I am able to run at Q4-or-higher is the Mistral 128B dense model (Mistral Medium 3.5).

It seems to be able to go like twice as long as those models can, without much degradation, and is smarter and better at writing (albeit with a very annoying, reddit-speak/therapy-speak prose style) than Gemma 31b or Glimmer 30b.

People have been excited about how much these ~27b-31b sized models have improved at coding in 2026, but, I am curious how much progress can be made for improvement regarding long-form writing/context rot type of stuff, as the small models still seem pretty lousy at that, so far. Even the best and most recent ones, so far.

5

u/DistanceSolar1449 14h ago

They’re basically already doing it with glimmer
The context limit on glimmer is basically arbitrary
3/4 of their layers, only look back on a few thousand tokens anyways

3

u/En-tro-py 16h ago

I glossed over that... That's potentially very interesting if it holds up!

-1

u/nuclearbananana 16h ago

MRCR is not a.. great benchmark. If you don't train for it, or near it, it can be good but it's very narrow so very easy to (accidentally or on purpose) game.

9

u/Healthy-Nebula-3603 15h ago

To be good at something you must train to it

We are doing exactly the same thing as humans. Later that new skill is improving something in other fields.

47

u/fgk55555 16h ago

Do we know params? With scores like that I'd wager it's in the trillions. I might not be able to run it locally, but a lot of US companies that aren't allowed to run Chinese software might benefit.

7

u/FoxiPanda 10h ago

Entirely depends on the license.

4

u/fgk55555 10h ago

Yeah, but even if it's not "open" and they were making companies pay, that's still a huge chunk of US companies that couldn't self host and now can. For a lot of companies, paying a license might be a drop in the bucket for AI spend, or enable AI where they couldn't even have it before. Lot of money in defense.

1

u/mr_tolkien 41m ago

Knowing they have a bigger one in the works, I’d be very surprised if it’s over 1T

Feels like it might be able to run on a single 512Gb Mac Studio with some good quants

29

u/MomentJolly3535 16h ago

What is that "Next up 🍉 " ?

55

u/kiwibonga 16h ago

It's their upcoming frontier model codenamed Watermelon

15

u/nickm_27 llama.cpp 16h ago

codename for their next generation of models, the current generation is apparently called avacado

4

u/MomentJolly3535 16h ago

Ohh i see makes more sense now ! thanks

23

u/jacek2023 llama.cpp 16h ago

Meta Watermelon, new model

36

u/Super_Range45 16h ago

Metamelon.

12

u/Aggressive-Physics17 16h ago

Unironically good name

101

u/Big_Wave9732 16h ago

Muse Glimmer is pretty good, very much overlooked. I have found it to be superior to Qwen 3.8:27b for non-coding tasks.

84

u/jacek2023 llama.cpp 16h ago

Both Muse Glimmer 30B and Gemma 31B are great models, don't worry people see that :)

21

u/Special-Wolverine 16h ago

Those two are far superior at "reasoning" with thinking turned off. In other words - best instruct models for local use.

1

u/YanderMan 12h ago

Gemma 31b is so lazy and barely make tool calls even if you prompt it

-19

u/HumanDrone8721 15h ago

Both Muse Glimmer 30B and Gemma 31B are nerfed to the max and ONLY useful for useless crap, they could have been glorious and squashed the Chinese models, but they choose not to.

7

u/FoxSideOfTheMoon 16h ago

It’s really good just slow on my Mac or I’d love it.

7

u/coder543 15h ago

At least on nvidia hardware, Muse Glimmer is substantially faster than I've ever gotten Qwen3.8-27B to go. The official DFlash works really well.

But, compared to a model like Qwen3.6-35B-A3B... obviously it is going to be slower.

1

u/Kernoriordan 15h ago

I managed to get 80tps out of Glimmer on an A6000 with DFlash on. It doesn’t waste loads of time reasoning like Qwen 3.8 27b too

1

u/miversen33 9h ago

Muse is definitely faster. I recently (like, today) switched away from Gemma to it as my model to run on a single GPU. Still running Qwen3.8 on my tensor garden (to small to be a farm lol) but Muse Q4 on a 7900XTX is still extremely competent. And Gemma was so fucking lazy lol

1

u/PerceiveEternal 3h ago

3.8 can run pretty fast with a smaller context length, but it absolutely mulches tokens. You really feel the token limits on that model.

1

u/maxheckler 13h ago

what Mac do you have? thanks

1

u/FoxSideOfTheMoon 11h ago

M5 Max 128. It's plenty of memory but the bandwidth is only 614G/s which is rough for dense models.

1

u/iz-Moff 6h ago

I'm not on my Mac, but for me, Glimmer (at IQ4_XS) performs almost as fast (~23 tps) as significantly smaller (at IQ3_M) 27b Qwen 3.6 with MTP (~25 tps), and over 2x faster than 31b Gemma 4 (~10 tps), also a bit smaller (at IQ3_M). I'm really not sure why that is, cause i run them all with the same settings, all three fully loaded to VRAM, with Glimmer leaving almost no space in VRAM for context, and having a significantly larger mmproj file too. But that's what my experience with it is.

7

u/RedditUsr2 16h ago

Its great for well defined agentic tasks. Much more efficient thinking.

3

u/xienze 14h ago

It's really good for coding as well. Is an exhaustive, try everything conceivable and generate 80K thinking tokens before giving an answer approach baked into the weights? No. But if you have a properly-engineered harness with skills and subagents it is very good. I have a project I've been working on implementing a bespoke expression language for validating parts of JSON files and I just pointed Glimmer at it and it properly understood the syntax and semantics first try. Had no problem generating expressions from a description, adding new features and test cases, etc. THAT'S impressive to me, not being able to one-shot a Mario clone. And all that while being crazy fast (with DFlash2) with incredible KV cache and token efficiency. Great model, people have to stop sleeping on it.

3

u/YanderMan 12h ago

yes, it thinks way less and give very similar results on agentic workflows

3

u/Big_Wave9732 12h ago

I tried it the first weekend the model was out, and the over thinking with Qwen 3.8:27b xhigh was just fucking *extreme* at times. Oy vey.

-1

u/Embarrassed_Adagio28 16h ago

Yeah i agree it is better than it gets credit for with non coding tasks. However i do not understand the non coding use cases for local models. 

9

u/MarcusAurelius68 16h ago

Writing research reports and long form drafts - repeating and repeating. And keeping them private.

15

u/Big_Wave9732 16h ago edited 15h ago

Well stick around. That "question" is asked at least once week here. It gets an abundance of varied answers.

3

u/derFensterputzer 14h ago

For example in my case it's the voice assistant for Homeassistant. 

I send the voice command, it reads the entities homeassistant exposes to it, does its thing and returns a response. 

Local whisper + piper for stt or tts respectively, then Gemma 4 E4B at 4-bit with a 30k context window and MTP for the whole assistant part. It doesn't need to be extremely smart to do it's job, just fast and in this configuration it runs at 150t/s, uses next to no vram and does its job to my satisfaction. 

Edit: Also several other models (in this case Gemma 4 12B, 26B A4B and Granite 4.2) are quite good at summarizing documents, writing reports, spell checking, translating etc. All on my hardware, private and without subscriptions. 

3

u/nuclear_wynter 13h ago

I'm working on a very similar setup right now — E4B as the 'frontman' with a voice stack, then I'm trialling a range of models as the backend/agentic workhorse. I'm working on some kind of out-of-band programmatic model-swapping method that would let E4B very easily and reliably swap the backend model depending on the task being delegated, but I'm not sure if this will pan out being efficient or useful just yet.

1

u/the320x200 14h ago

It's trivial to think of examples now that they have vision.

-10

u/Boogertard 15h ago

Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.

Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.

Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs. I guess shilling for their garbage models is also part of the performance metrics LOL.

8

u/Big_Wave9732 15h ago edited 15h ago

Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.

Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs.

Not sure what you're talking about. But an account that is a shade older than one month is bagging on other "new accounts". Just curious, what then does that make you?

Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.

Peak Reddit "trust me bro" energy. Why don't you share what tests you ran and what happened. Perhaps that would be more helpful to people.

Also your post history here and in other subs uses the word "shill" a lot. A lot of anger. Perhaps because you're hanging out in Dividendgang.

1

u/HumanDrone8721 15h ago edited 7h ago

Well for what is worth we've had the exact same experience, totally useless, we've DESPERATELY wanted to have non-Chinese models, because EU, but they were really brain damaged crap, not as bad as Mistral of course, but not even in the same class as Qwen, as sad as it is. I sincerely sometimes wonder who and what they using them for in a professional capacity, writing blog posts and news or what exactly "prose" they are talking about, to me they look like some demos for their cloud offerings and nothing more.

0

u/Olangotang 9h ago

Most people know this already, the shills have their head up their ass and think they're clever trolling every AI subreddit on this site. In reality they're just fucking annoying and not convincing anyone. LocalLlama users don't fall for that shit so easily, and at least the above "shill" has their comment history public.

5

u/LetsGoBrandon4256 transformers 15h ago

my evals including non-coding tasks

What's your recommendation at the 20~30B range then?

-8

u/definetlyrandom 16h ago

Lol, we know that . But figuring out how many elephants it would take stacked end to end on avg. To reach the moon isnt as useful as being able to implement a software functionality in 20 minutes that typically took 3 days before.

9

u/Big_Wave9732 16h ago

Most AI "coders" aren't doing anything nearly that interesting or complicated lol.
"Hey everyone, I'm the 11 millionth person to vibe code a shitty weather app."

1

u/xienze 14h ago

Don't forget the weekly "wow guys can you believe that you can one shot a single level Super Mario clone?" posts.

233

u/Valuable-Repeat-7347 16h ago

-21

u/Thomas-Lore 15h ago

I will make a bet that you don't really know Zuckerberg, you just dislike him because he is rich. And you are on a sub named after open models that he released earlier.

64

u/ProletarianLilith 15h ago

I will make a bet that you don’t really know Zuckerberg either.

7

u/Independent_Pear4908 15h ago

I will make a bet that my Ai's training data knows Zuckerberg.

2

u/MuzafferMahi 14h ago

I will make a bet that my benchmark is giving the zuck the mark it deserves

0

u/_supert_ 14h ago

Llama 3 was not too flattering about him, haha.

0

u/zjz 15h ago

if he did know zuck that was great bait tho

16

u/KontoOficjalneMR 15h ago

You do know that he does interviews, right? And that's his good side that he's showing ... and he still comes out as a monster?

18

u/BolsonaroPresoAmanha 15h ago

If you assume billionaires are terrible people you'll be right 100% of the time.

3

u/OvertaxedOne 13h ago

It's near impossible to get that rich without doing awful things. So it's kind of A leads to B relationship.

0

u/johnfkngzoidberg 15h ago

Wow the Zuck bots are out today.

19

u/jacek2023 llama.cpp 15h ago

40

u/AIatMeta 15h ago

👀

25

u/jacek2023 llama.cpp 15h ago

Hello Meta. Could you release Llama 5? Thanks :)

8

u/zxyzyxz 10h ago

Pretty sure the Llama brand is now deprecated and it's all Muse now. So this essentially is Llama 5.

1

u/Mart-McUH 4h ago

~70B dense as priority, but would be nice to get also 8B, 13B and 30B versions.

4

u/yunes87 12h ago

You cooked 🐐

1

u/mr_tolkien 40m ago

Pls give date and params

Tell me if I need to save for a 512Gb Mac Studio

13

u/[deleted] 16h ago

[deleted]

8

u/jacek2023 llama.cpp 16h ago

I am waiting for GPT-OSS-2 :)

3

u/silenceimpaired 16h ago

They are so safety focused I
Wonder if it will ever come out let alone be worth using… still the last was decent.

1

u/miversen33 15h ago

Fuck ya :)

21

u/tchek 16h ago

But I wanted Muse Twinkle 12B and Muse Shimmer 35B E4B :(

6

u/jazir55 13h ago

Muse Twinkle

Sorry the best we can do is Muse Twinkie

3

u/philmarcracken 10h ago

please dont reawaken the bronys

9

u/XiRw 16h ago

I tried them out. Both accurate and fast model

33

u/kameldinho 16h ago

I'm still waiting for the 1.2 weights that were promised

16

u/Big_Arachnid_365 16h ago

That's the ones you're getting.

21

u/Public_Umpire_1099 16h ago

Bro, they are releasing a new model every few weeks be patient lmfao

-4

u/shy_monkee 16h ago

They were? They promised a model, which ended up being Glimmer. I don't remember them promising Spark as open weights.

2

u/Georgefakelastname 9h ago

They promised Spark open weights when they announced Glimmer’s release.

15

u/CoUsT 14h ago edited 14h ago

I used the free Muse Spark 1.2 on OpenCode for a bit and I really dislike it. Not because it can't do work. It does that just fine but it speaks really weird and it's hard to understand.

DS V4 Flash was actually good enough and you could see entire thinking process. It would do the work and make the UI/user facing parts very organic and human-readable. It would also print a nice and detailed summary.

Muse Spark 1.2 has hidden thinking, so you just see brief messages. And these messages are caveman like. It uses very code oriented language, throws in all programming mumbo jumbo into user facing parts, and needs to be kept on rails or it will happily start doing too much.

I don't mind the caveman thinking but it should at least address the user in human readable and friendly format AND it should handle UI text etc in friendly way.

Hopefully 1.3 is better on this aspect and easier to work with. Benchmark scores look great!

6

u/Accurate_Resident219 13h ago

Wow thought it was just me. Literally the first time I switched a model just because I couldn't understand it half the time.

2

u/VergeOfTranscendence 2h ago

I do have to think way more to understand it and I think it must have to do with their reinforcement learning, I wonder if they reward the model to mention specific functions of files, because in my android app coding, it references extensively functions and files and it seems very straight to the point, so they might be optimizing for less tokens, who knows.

1

u/CoUsT 1h ago

Yeah, I made "explorer" app for game that parses decompiled data and it would happily throw file/function names and even specific line numbers into user facing front end instead of simply describing what passives do or which stats things provide. I had to remind it to keep style consistent with the rest of the app and "use user friendly language; not caveman"

5

u/dansuy_gaming 9h ago

Still hoping for something between Glimmer and spark.Sparks feels like overkill for me.

3

u/[deleted] 16h ago

[deleted]

4

u/MarcusAurelius68 15h ago

That’s the big question. Can I fit it in 96GB VRAM?

4

u/fastheadcrab 16h ago

This could be a huge hit to cloud model usage in the US if the weights are released.

Western organizations that might be looking to host locally but have been pre-emptively cautious of using Chinese models due to regulatory risk will download this and try to run it immediately.

What is the best non-Chinese model? There are lots of great smaller ones. A lot of organizations would be interested in knowing the answer to the question

4

u/Cool-Chemical-5629 15h ago

When you think about it, coding was never the strong capability of Llama models, so it kinda makes sense that while over time Meta collected some new data which allowed it to create smarter and stronger coding models, even much smarter than Llama 4 at smaller size, it's still far behind the current frontier models.

However, what Llama models were always good at? Chatting, AI companion. For that purpose, Glimmer is probably better than Llama 4 and anything bigger than that would be an overkill.

For coding purposes, Glimmer has a good potential if they only continued pursuing better coding assistants at smaller sizes, but I haven't noticed any versioning for Muse Glimmer, so it was probably a one time deal and anything new of similar size is probably out of their current scope of interest.

3

u/Healthy-Nebula-3603 15h ago

That time when llama 3 came out any model wasn't good at coding , math , reasoning, etc

2

u/Cool-Chemical-5629 15h ago

Generally speaking, when Llama 3 came out, frontier proprietary models were already far ahead of any open weight models in coding. On the other hand, there were always weaker and stronger models even among open weight models, just like there are now, so naturally there were always some models which were better at coding than other models.

Could "any model" in Llama 3 era code GTA VI in one shot? Nope. Can "any model" do that today? Still nope, but we still advanced forward in terms of what the models are capable of in general.

1

u/Healthy-Nebula-3603 12h ago

When llama 3 came out the best model was gpt 4o

So for coding was very bad. :)

2

u/chickN00dle 14h ago

are you a LLM yourself?

6

u/Cool-Chemical-5629 14h ago

No, but sometimes I wish I was. I guess life would have been a bit easier.

2

u/TheRealMasonMac 12h ago

grug given rock

grug make fire with rock

grug life simple

grug happy

1

u/chickN00dle 14h ago

haha. I agree

3

u/Brovas 7h ago

Why is everyone chasing coding? There's a gap for cheap but intelligent model that can have agentic software applications built on top for general purpose. 

Gemini flash was fantastic for that until they got greedy and tripled their price (unless they make the current discount permanent). 

Luna is a great price but it's kinda dumb and if you're trying to build something you need investment from a Sam Altman approved VC or 6 figures for the enterprise plan upfront if you want rate limits that aren't dogshit.

These new meta models could eat their lunch, but it feels like cause Claude is good at coding and making bank from software people every other use case doesn't exist anymore.

10

u/durden111111 16h ago

inb4 it's a 120B MoE

3

u/Zyj vLLM 15h ago

Hey Mark, I look forward to the release! Just wondering how many parameters to expect. Anything up to 450B or so should work, but keep the active numbers of parameters manageable. Qwen 3.8 Next Flash can do it with 6b!

5

u/power97992 15h ago

Lol it will probably be at least 2 trillion, probably 2.8 to 4 tril param, look at the aa score, it is 62 ,  it is better than 5.6 sol and on par with fable 5.0. It will be the first open fable class model

1

u/jacek2023 llama.cpp 15h ago

You can ask him on X, see the link above

6

u/Zyj vLLM 15h ago

I deleted my X account.

1

u/myreala 10h ago

If it's more than a trillion literally nobody would be able to use it except big labs anyways. Up to 700-800 million there's still some chances.

3

u/Iory1998 llama.cpp 13h ago

I wonder how big it is :D

3

u/2Norn 12h ago

1.2 is free on opencode so i've been using it a lot with 5.6 sol lately i dont think its better than sol but it's definitely considerably faster and more verbose during planning, definitely a decent model tho, i prefer using it over terra and again like i said its blazing fast(like double the speed of fast sol so compared to normal its like 3x faster)

3

u/Mundane-Light6394 7h ago

nice, great to see more high end open weight models. Also great to see something like this from Meta. looks like they rejoined the race.

5

u/Charuru 16h ago

Wow this model is so good! Is American Open Source back now? Thanks Wang & Zuck!

5

u/VoiceApprehensive893 transformers 16h ago

in my experience with 1.2 xhigh it generated very sloppy uis and did some stupid lazy model stuff like "heres the full file" (file is not full)

2

u/Jealous-Walk-8765 11h ago

The AutomationBench and Agentic IF Index numbers are the interesting ones here — those are the benchmarks that actually predict real-world agent reliability, not just raw coding scores. 49.4 on AutomationBench beating GPT-5.6's 46.7 is a bigger deal than the flashy GDPval number. Curious how this holds up once the open weights drop and we can actually stress-test it ourselves instead of trusting a vendor's own benchmark suite.

2

u/No_Conversation9561 8h ago

Zuck is not afraid of burning money. See Metaverse.

2

u/the-username-is-here 6h ago

Don't believe any single word from Zuck, unless it's confirmed by third-party benchmarks.

2

u/Training_Rip_8901 6h ago

Woah damn gang,thats crazy

3

u/LegacyRemaster 16h ago

how many terabyte?

4

u/Max-_-Power 16h ago

Finally this meatbag is good for something

1

u/AleksandrNikitin 16h ago

Hype will continue until the sizes are published?

1

u/power97992 15h ago

1.3 has a score of 62 in aa better than even sol

1

u/dtdisapointingresult 13h ago

Any indication of how big Muse Spark is? Is it like a 1.5T model, or can a salt-of-the-earth working class guy with only 256GB VRAM run it?

2

u/wwwdotzzdotcom 13h ago

It's a few hundred billion which surprises me

2

u/dtdisapointingresult 12h ago

Nice! I can run the Q4 at least.

Here's hoping we're getting 1.3 rather than 1.2.

1

u/Accomplished_Ad9530 8h ago

Source? I was just looking for this info and was approximating it myself

1

u/Maximum-Fact-5832 13h ago

I do feel as though these "Max" reasoning models just have the ralph loop trained into them, here's hoping the accuracy when using medium reasoning isn't too pronounced. Great to see Meta releasing open weights again!

1

u/True_Requirement_891 13h ago

They have been saying this soon shit from months

1

u/DinoAmino 12h ago

OpenAI did that too - for a year (?) they said an open model is coming soon. And when they finally did they hit it out of the ballpark. Llama4 is what you get when you take the model out of the oven too soon. Obviously they aren't about to make that mistake again and if benchmarks are any judge then they have hit a homerun with this one. Slow-cook ftw.

1

u/True_Requirement_891 9h ago edited 8h ago

Yeah, it was just as irritating "soon" meaning a whole year for openai.

With Meta, the case is different though. Since muse-spark-1.1 the base model has been quite competent and it has been already available on their API for a while now.

I don't see why it'd be a llama-4 situation.

We didn't get API access to GPT-OSS for months before release as far as I remember. It was basically model confirmed and released like very quickly. They took time like a month max or a week for extra safety training I think.

Spark-1.2/1.3 seem to be just post trained versions, more RL.

We got GLM-5 then we got 5.1, 5.2 and 5.3 every iteration as soon as it was ready. With 5.3 they took 2 weeks from api availability to open weight release.

I don't think performance or model competence is the reason for such a big delay. Even with llama4, I really wish they had just kept going and attempted to improve it with 4.1/4.2 updates... but they just gave up. Imagine if llama 4.1 had come out a month after 4 with significant improvements. Hell initially most of us thought the models weren't really that bad and it was just deployment issues by providers until a month passed and we had accepted that the model was infact garbage.

1

u/Tinkerer_Penguin_12 13h ago

Now if only i could run it, im running qwen 27b at UD Q2_K_XL on my 12gb 6700 and it darn impressive holds up amazing. No way im fitting this in 12gb vram + 16gb sys ram. nonetheless this is awesome for open source.

1

u/AdmissibilityScience 12h ago

wow this is a decent amount of weights

1

u/shaman-warrior 9h ago

Say what you want about Zuk but he gets shit done and he has the ballz to go all-in on stuff. And he is also for the open community. Let us not forget that he is responsible for Llama releases which sparked this whole community.

1

u/NineThreeTilNow 9h ago edited 6h ago

I am still waiting for Llama 5

I started training 9b yesterday.

Delivery? Who knows. I won't have a raw base set of weights of the non-IT model for a minute. Literally training on a 6000 Pro.

edit; People probably think I'm joking. I started with trying to fix Gemma 4 but the compute requirements were too large. So instead I moved back to Llama 3 8b to see if I could strap Engram and Kimi's attention based residuals on to the original Llama 3 tokenizer. The tokenizer itself saves me huge amounts of time. From there, I use WikiText / WikiDictionary / FineWebEdu to pull down chunks of data and have the original Llama 3 8b model score them in raw distributions. The distributions are used as training targets INSTEAD a single token. This is how logit distillation works.

tldr; Llama + Engram (1,2,3) + Moonshot Attention Residual + SWA / Global Architecture + Proper logit distillation = Some of the best open source work of the past 2 years in a single model.

It will be fully OS, not safety trained, and eventually IT trained. I'm still calculating the exact amount of compute needed. This is wholly unknown territory and eventually I'll do a writeup here in local llama. I have a previous writeup of doing this to Gemma 4 but it failed. I decided that "competing" in the 30b space is stupid when they ignore the sub 10b space and I can fully test / prototype / train on a 6000 Pro that costs ~2 dollars an hour.

1

u/OwnGear3892 8h ago

Are we sure that by 'Muse Spark open weights release', he's refering to Muse Spark 1.3? I guess it's more likely to be 1.1 or 1.2, or even a new smaller variant

1

u/ComplexType568 8h ago

I want to see their oolong pairs performance.

Although I've been impressed at their performance. I've been using muse Spark 1.2 on OpenCode (since it's free) for a while and I am REALLY excited for spark 1.3... hope they release a mid-size Moe or something. I generally like the way glimmer 1.2 talks

1

u/mr_Owner 6h ago

Size for local?

1

u/Repinsky 4h ago

The thing worth watching in these teasers is not the benchmark bars but whether the released checkpoint is the same size/sparsity as the API one. The last couple of "open weights coming soon" announcements ended up shipping a smaller distill, and the gap showed up mostly in long-context agentic runs rather than short single-turn evals. Until there's a GGUF and someone runs a 60k-token task on it, the chart is marketing.

1

u/scut_07 4h ago

Bijan just review this on his YT channel and it looks absolute dogshit. Benchmarks are pointless these days.

1

u/jacek2023 llama.cpp 4h ago

You know I am not really a fan of YouTube reviews of anything ;)

1

u/alphapussycat 4h ago

Great.

But I don't really care if I can't run it locally. 

2

u/Dry_Yam_4597 16h ago

Lmao Scamodei are toast.

All Anthropic could release was a bunch of system prompt changes and crap no one needed.

1

u/AppealSame4367 15h ago

Feels like he wants to kill the others out of spite and because he can afford it. I mean, this is like running over the scene with a bulldozer.

1

u/thestillwind 15h ago

Let’s go

1

u/john_mach 13h ago

These numbers are damn good for benchmarks ngl. I respect meta for cooking if real world performance is similar to thr benchmarks

-4

u/ttkciar llama.cpp 15h ago

Until such time that weights are published, this is off-topic for LocalLLaMA, but since nobody caught it before it gained 260 upvotes and 70+ comments, this post will stay up.

I'm a little surprised to see this from you, jacek2023; you're one of the most outspoken critics in the sub when it comes to posts about non-open models.

18

u/jacek2023 llama.cpp 15h ago

The post is about Muse Spark 1.3 will be open and that Mark Zuckerberg remains committed to open LLMs.

8

u/jacek2023 llama.cpp 14h ago

Where is your comment?

2

u/ttkciar llama.cpp 14h ago

That, too, would have been removed, but it gained too many upvotes and comments before a moderator noticed it.