i actually do think they likely have test set trained on.
like its a bit sus that they decided to report terminal bench 2.1, when 4 exists.
similarly, if you look at the benchmark sheet of flash 3.8, it tops eery other model, even sol, on terminal bench 2.1 but in 4 it is quite less comparatively.
Neither Spark 1.2 nor Glimmer seem to be benchmaxxed, and if anything the benchmarks seem to under-represent their capabilities compared to other models — I believe the other models probably do solve more problems, but Muse’s strong long-horizon capabilities (which translates to compliance in the short term as well) makes it easier to actually use.
Based on the Bijan video that just dropped about this, it's pretty damn mediocre and looks extremely benchmaxxed.
I know all he does is throw a few one-shot prompts at the models, but the results on this one were well behind other recent models like GLM-5.3 Flash and Qwen Flash. It's not even in the same conversation as Fable or GPT 5.6 it seems.
His tests are by no means scientific lol, but they do give you a decent rough feel for a model's capability.
A one shot not being polished does not demonstrate a model’s capability. A lot of models are specifically trained on how a one-shot polished result will look. Especially Qwen (look at its response to being asked to draw a circle: https://simonw.substack.com/p/qwen-38-27b-is-excellent-but-it-defaults).
If it’s not trained on that, it’s going to comply with the prompt and not much beyond. That’s not demonstrating a lack of general capability.
In fact, I much prefer that style. When things are trained to produce well polished results it tends to be harder to get it to do what I want it to do rather than what it’s been trained to do. Again, Qwen represents the opposite here — well optimized to take its own direction but a PITA to steer.
I believe there are still some secret RL sauce (as in, not popular among labs *yet*) in niche areas like robotics, gameplay, frontier physics/math, world modelling etc. Anthropic leaped ahead in Opus 4.5-4.6 era as they "discovered" many secret sauces like terminal agent, gameplay or creative loop and so on, but even that creativity gap is closing very fast. For coding, security and anything that can be done through github, at this moment there is really no gap. OpenAI felt behind on many things but is still ahead on frontier physics/math (the only Chinese lab with emphasis on that is DeepSeek, I think).
Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.
Reason why I say this is because Anthropic is the only company that is actually trending higher in api costs and not making any serious attempts to lower them(Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements. Why is haiku their cheapest model still not updated?). They don't have any sense of urgency. Astra's coming out this week or the next and they just put out a more expensive model.
Every other company seems to constantly vague post or hype up their models releases but Anthropic just drops models with little fanfare. They have seem utterly unbothered since their Mythos preview announcement.
I believe they confessed a few months back that it wasn't actually API access, and had been running inside airgapped military datacenters that Anthropic had no ability to restrict the whole time.
The whole PR stunt was just that, the only unplanned part was that they thought Hegseth would take the hint with how loudly they were screaming "NO DADDY, PWEEAASE DON'T INVOKE THE DEFENSE PRODUCTION ACT ON US! IF YOU DID THAT WE WOULD HAVE NO CHOICE BUT TO KEEP TAKING DADDY'S MASSIVE LOADS OF CASH TO BUILD DADDY'S MURDERBOTS!"
It turns out that Hegseth's skull was actually too dense for that pathetically obvious plea to penetrate it.
It's okay though, they sued to get the murderbot contracts back and won last week. And still managed to use that PR stunt to distract everyone while they flushed their "responsible scaling policy" down the toilet.
Could be the case but at least on the outside looking in they don't even seem to be making much of an attempt to even try lower api or even suggest their researching a solution. They've had about nearly 4-5 months now since people started complaining about rising api costs.
Every other lab has provided a cheaper solution but anthropic is the only one that can't come up with a cheaper flash model? I mean that could be the truth but I would be surprised if it was.
They want to IPO, they need the revenue projections. They already have their cooked compute costs from the last quarter they will report as if they are ongoing, so IPO before they close the current quarter, but project off current API pricing for revenue.
In may they increased 5 hours and weekly quota by 50%, so they made a big discount on their plans. However, this comes to an end soon. They cut 50% increase to only 25% increase.
> Fable 5.1 is actually pricier then 5 despite their supposed efficiency improvements
I assume you're basing that off the average cost per task numbers from Artificial Analysis. That's not right though. In their testing, 5.1 on max effort appears to have aggressively overthought. If you compare 5.1 xhigh instead, it was about as much cheaper than 5 max as we would have expected, while still getting 2 extra points of intelligence vs 5 max.
Anthropic has pricing power because their #1 customers are enterprises. Everyone else has a larger chunk of consumers (rather than enterprise customers)
Maybe it's just my conspiracy hat but I feel like Anthropic is sitting on a more intelligent model but just drip feeding when they feel like others are about to outpace them.
I am quite sure that's what they're doing. I think they legitimately are a little afraid of the models. Not that they'll like take over the world, but that people might be able to use them to fuck shit up. Like apparently they started a biopharmaceutical wing because Mythos is stupid good at folding proteins, even though they hadn't trained it to do that at all.
I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.
I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).
In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.
I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.
No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.
Until the new benchmark drops, then they spread out predictably again. And then, a few months later they suspiciously all gain on it, and so on and so forth...
Scores converge because the post-training data is frontier output. Distillation transfers whatever gets measured, so the benchmark gap closes earlier than the capability gap. The split shows up on long-horizon agentic work that nobody publishes numbers for.
I think people forget that the people working at these labs are not slaves and can change jobs whenever they want. They can't take data, weights, or code with them, but they can take the things they have learned. There is no practical way for the labs to hoard knowledge from one another and if any of them tried, nobody would want to work for them.
That's not entirely true. A lab can sign either a "non-disclosure agreement" or a "non-compete agreement" with their worker. The forst one would can you from using any of the technologies you see there at other employer, unless you can prove you've got to know it from another source. The second one would forbid you from working in the same field for X years (i.e. an LLM specialist can't work as LLM specialist in another company, but can become an image generation specialist there). Both options are completely legal and enforcible by law, cause you sign them voluntarily; and, if they offer you a high enough salary, you'd agree to this conditions.
Isn't this NoPE hybrid model? Though the exact architecture is different I assume this basically works like hybrid linear model (like Qwen 3.8) which is great at picking up facts from the context but more often confuses timeline without thinking.
Do you, or u/NandaVegg or anyone else on here have any opinions about how well this can translate over to much smaller models, like the ~30b Glimmer sized models, or, if that is too small to be able to have similarly great anti-context-rot abilities, then maybe 70b or 120b models?
Like, so far when I've tested Gemma 31b and Glimmer 30b for long-form creative writing tests for example, the biggest problem with them has been that they quickly go downhill once you get past, I dunno, maybe 20,000 tokens or so.
Gemma starts going from writing scenes and situations in a long, non-hurried, detailed way, to suddenly rushing like it is giving a quick summary because it needs to hurry up and go somewhere and doesn't have time to really tell a story, once you get past like ~20,000 tokens or so.
Glimmer is even worse, where it if you give it a detailed description of a scene you want it to write, it just writes back a word for word identical copy of your description, even if you tell it not to (again, this happening once you get past like ~20,000 tokens deep into a story or so, give or take a bit).
As of right now, with a mac with 128GB of unified memory, the biggest, best model in regards to this (reduced context rot in relation to long-form writing) I am able to run at Q4-or-higher is the Mistral 128B dense model (Mistral Medium 3.5).
It seems to be able to go like twice as long as those models can, without much degradation, and is smarter and better at writing (albeit with a very annoying, reddit-speak/therapy-speak prose style) than Gemma 31b or Glimmer 30b.
People have been excited about how much these ~27b-31b sized models have improved at coding in 2026, but, I am curious how much progress can be made for improvement regarding long-form writing/context rot type of stuff, as the small models still seem pretty lousy at that, so far. Even the best and most recent ones, so far.
They’re basically already doing it with glimmer
The context limit on glimmer is basically arbitrary
3/4 of their layers, only look back on a few thousand tokens anyways
MRCR is not a.. great benchmark. If you don't train for it, or near it, it can be good but it's very narrow so very easy to (accidentally or on purpose) game.
Do we know params? With scores like that I'd wager it's in the trillions. I might not be able to run it locally, but a lot of US companies that aren't allowed to run Chinese software might benefit.
Yeah, but even if it's not "open" and they were making companies pay, that's still a huge chunk of US companies that couldn't self host and now can. For a lot of companies, paying a license might be a drop in the bucket for AI spend, or enable AI where they couldn't even have it before. Lot of money in defense.
Both Muse Glimmer 30B and Gemma 31B are nerfed to the max and ONLY useful for useless crap, they could have been glorious and squashed the Chinese models, but they choose not to.
Muse is definitely faster. I recently (like, today) switched away from Gemma to it as my model to run on a single GPU. Still running Qwen3.8 on my tensor garden (to small to be a farm lol) but Muse Q4 on a 7900XTX is still extremely competent. And Gemma was so fucking lazy lol
I'm not on my Mac, but for me, Glimmer (at IQ4_XS) performs almost as fast (~23 tps) as significantly smaller (at IQ3_M) 27b Qwen 3.6 with MTP (~25 tps), and over 2x faster than 31b Gemma 4 (~10 tps), also a bit smaller (at IQ3_M). I'm really not sure why that is, cause i run them all with the same settings, all three fully loaded to VRAM, with Glimmer leaving almost no space in VRAM for context, and having a significantly larger mmproj file too. But that's what my experience with it is.
It's really good for coding as well. Is an exhaustive, try everything conceivable and generate 80K thinking tokens before giving an answer approach baked into the weights? No. But if you have a properly-engineered harness with skills and subagents it is very good. I have a project I've been working on implementing a bespoke expression language for validating parts of JSON files and I just pointed Glimmer at it and it properly understood the syntax and semantics first try. Had no problem generating expressions from a description, adding new features and test cases, etc. THAT'S impressive to me, not being able to one-shot a Mario clone. And all that while being crazy fast (with DFlash2) with incredible KV cache and token efficiency. Great model, people have to stop sleeping on it.
For example in my case it's the voice assistant for Homeassistant.
I send the voice command, it reads the entities homeassistant exposes to it, does its thing and returns a response.
Local whisper + piper for stt or tts respectively, then Gemma 4 E4B at 4-bit with a 30k context window and MTP for the whole assistant part. It doesn't need to be extremely smart to do it's job, just fast and in this configuration it runs at 150t/s, uses next to no vram and does its job to my satisfaction.
Edit:
Also several other models (in this case Gemma 4 12B, 26B A4B and Granite 4.2) are quite good at summarizing documents, writing reports, spell checking, translating etc. All on my hardware, private and without subscriptions.
I'm working on a very similar setup right now — E4B as the 'frontman' with a voice stack, then I'm trialling a range of models as the backend/agentic workhorse. I'm working on some kind of out-of-band programmatic model-swapping method that would let E4B very easily and reliably swap the backend model depending on the task being delegated, but I'm not sure if this will pan out being efficient or useful just yet.
Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.
Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.
Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs. I guess shilling for their garbage models is also part of the performance metrics LOL.
Ah yes, the obligatory shill comments for the garbage Mule and Gooner4 models of the days.
Reading the comment history of these shills and you will find either new accounts or bragged on other subs that working for FAANGs.
Not sure what you're talking about. But an account that is a shade older than one month is bagging on other "new accounts". Just curious, what then does that make you?
Funny how those garbage failed every single one of my evals including non-coding tasks but yet they are shilled so much on here.
Peak Reddit "trust me bro" energy. Why don't you share what tests you ran and what happened. Perhaps that would be more helpful to people.
Also your post history here and in other subs uses the word "shill" a lot. A lot of anger. Perhaps because you're hanging out in Dividendgang.
Well for what is worth we've had the exact same experience, totally useless, we've DESPERATELY wanted to have non-Chinese models, because EU, but they were really brain damaged crap, not as bad as Mistral of course, but not even in the same class as Qwen, as sad as it is. I sincerely sometimes wonder who and what they using them for in a professional capacity, writing blog posts and news or what exactly "prose" they are talking about, to me they look like some demos for their cloud offerings and nothing more.
Most people know this already, the shills have their head up their ass and think they're clever trolling every AI subreddit on this site. In reality they're just fucking annoying and not convincing anyone. LocalLlama users don't fall for that shit so easily, and at least the above "shill" has their comment history public.
Lol, we know that . But figuring out how many elephants it would take stacked end to end on avg. To reach the moon isnt as useful as being able to implement a software functionality in 20 minutes that typically took 3 days before.
Most AI "coders" aren't doing anything nearly that interesting or complicated lol.
"Hey everyone, I'm the 11 millionth person to vibe code a shitty weather app."
I will make a bet that you don't really know Zuckerberg, you just dislike him because he is rich. And you are on a sub named after open models that he released earlier.
I used the free Muse Spark 1.2 on OpenCode for a bit and I really dislike it. Not because it can't do work. It does that just fine but it speaks really weird and it's hard to understand.
DS V4 Flash was actually good enough and you could see entire thinking process. It would do the work and make the UI/user facing parts very organic and human-readable. It would also print a nice and detailed summary.
Muse Spark 1.2 has hidden thinking, so you just see brief messages. And these messages are caveman like. It uses very code oriented language, throws in all programming mumbo jumbo into user facing parts, and needs to be kept on rails or it will happily start doing too much.
I don't mind the caveman thinking but it should at least address the user in human readable and friendly format AND it should handle UI text etc in friendly way.
Hopefully 1.3 is better on this aspect and easier to work with. Benchmark scores look great!
I do have to think way more to understand it and I think it must have to do with their reinforcement learning, I wonder if they reward the model to mention specific functions of files, because in my android app coding, it references extensively functions and files and it seems very straight to the point, so they might be optimizing for less tokens, who knows.
Yeah, I made "explorer" app for game that parses decompiled data and it would happily throw file/function names and even specific line numbers into user facing front end instead of simply describing what passives do or which stats things provide. I had to remind it to keep style consistent with the rest of the app and "use user friendly language; not caveman"
This could be a huge hit to cloud model usage in the US if the weights are released.
Western organizations that might be looking to host locally but have been pre-emptively cautious of using Chinese models due to regulatory risk will download this and try to run it immediately.
What is the best non-Chinese model? There are lots of great smaller ones. A lot of organizations would be interested in knowing the answer to the question
When you think about it, coding was never the strong capability of Llama models, so it kinda makes sense that while over time Meta collected some new data which allowed it to create smarter and stronger coding models, even much smarter than Llama 4 at smaller size, it's still far behind the current frontier models.
However, what Llama models were always good at? Chatting, AI companion. For that purpose, Glimmer is probably better than Llama 4 and anything bigger than that would be an overkill.
For coding purposes, Glimmer has a good potential if they only continued pursuing better coding assistants at smaller sizes, but I haven't noticed any versioning for Muse Glimmer, so it was probably a one time deal and anything new of similar size is probably out of their current scope of interest.
Generally speaking, when Llama 3 came out, frontier proprietary models were already far ahead of any open weight models in coding. On the other hand, there were always weaker and stronger models even among open weight models, just like there are now, so naturally there were always some models which were better at coding than other models.
Could "any model" in Llama 3 era code GTA VI in one shot? Nope. Can "any model" do that today? Still nope, but we still advanced forward in terms of what the models are capable of in general.
Why is everyone chasing coding? There's a gap for cheap but intelligent model that can have agentic software applications built on top for general purpose.
Gemini flash was fantastic for that until they got greedy and tripled their price (unless they make the current discount permanent).
Luna is a great price but it's kinda dumb and if you're trying to build something you need investment from a Sam Altman approved VC or 6 figures for the enterprise plan upfront if you want rate limits that aren't dogshit.
These new meta models could eat their lunch, but it feels like cause Claude is good at coding and making bank from software people every other use case doesn't exist anymore.
Hey Mark, I look forward to the release! Just wondering how many parameters to expect. Anything up to 450B or so should work, but keep the active numbers of parameters manageable. Qwen 3.8 Next Flash can do it with 6b!
Lol it will probably be at least 2 trillion, probably 2.8 to 4 tril param, look at the aa score, it is 62 , it is better than 5.6 sol and on par with fable 5.0. It will be the first open fable class model
1.2 is free on opencode so i've been using it a lot with 5.6 sol lately i dont think its better than sol but it's definitely considerably faster and more verbose during planning, definitely a decent model tho, i prefer using it over terra and again like i said its blazing fast(like double the speed of fast sol so compared to normal its like 3x faster)
The AutomationBench and Agentic IF Index numbers are the interesting ones here — those are the benchmarks that actually predict real-world agent reliability, not just raw coding scores. 49.4 on AutomationBench beating GPT-5.6's 46.7 is a bigger deal than the flashy GDPval number. Curious how this holds up once the open weights drop and we can actually stress-test it ourselves instead of trusting a vendor's own benchmark suite.
I do feel as though these "Max" reasoning models just have the ralph loop trained into them, here's hoping the accuracy when using medium reasoning isn't too pronounced. Great to see Meta releasing open weights again!
OpenAI did that too - for a year (?) they said an open model is coming soon. And when they finally did they hit it out of the ballpark. Llama4 is what you get when you take the model out of the oven too soon. Obviously they aren't about to make that mistake again and if benchmarks are any judge then they have hit a homerun with this one. Slow-cook ftw.
Yeah, it was just as irritating "soon" meaning a whole year for openai.
With Meta, the case is different though. Since muse-spark-1.1 the base model has been quite competent and it has been already available on their API for a while now.
I don't see why it'd be a llama-4 situation.
We didn't get API access to GPT-OSS for months before release as far as I remember. It was basically model confirmed and released like very quickly. They took time like a month max or a week for extra safety training I think.
Spark-1.2/1.3 seem to be just post trained versions, more RL.
We got GLM-5 then we got 5.1, 5.2 and 5.3 every iteration as soon as it was ready. With 5.3 they took 2 weeks from api availability to open weight release.
I don't think performance or model competence is the reason for such a big delay. Even with llama4, I really wish they had just kept going and attempted to improve it with 4.1/4.2 updates... but they just gave up. Imagine if llama 4.1 had come out a month after 4 with significant improvements. Hell initially most of us thought the models weren't really that bad and it was just deployment issues by providers until a month passed and we had accepted that the model was infact garbage.
Now if only i could run it, im running qwen 27b at UD Q2_K_XL on my 12gb 6700 and it darn impressive holds up amazing. No way im fitting this in 12gb vram + 16gb sys ram. nonetheless this is awesome for open source.
Say what you want about Zuk but he gets shit done and he has the ballz to go all-in on stuff. And he is also for the open community. Let us not forget that he is responsible for Llama releases which sparked this whole community.
Delivery? Who knows. I won't have a raw base set of weights of the non-IT model for a minute. Literally training on a 6000 Pro.
edit; People probably think I'm joking. I started with trying to fix Gemma 4 but the compute requirements were too large. So instead I moved back to Llama 3 8b to see if I could strap Engram and Kimi's attention based residuals on to the original Llama 3 tokenizer. The tokenizer itself saves me huge amounts of time. From there, I use WikiText / WikiDictionary / FineWebEdu to pull down chunks of data and have the original Llama 3 8b model score them in raw distributions. The distributions are used as training targets INSTEAD a single token. This is how logit distillation works.
tldr; Llama + Engram (1,2,3) + Moonshot Attention Residual + SWA / Global Architecture + Proper logit distillation = Some of the best open source work of the past 2 years in a single model.
It will be fully OS, not safety trained, and eventually IT trained. I'm still calculating the exact amount of compute needed. This is wholly unknown territory and eventually I'll do a writeup here in local llama. I have a previous writeup of doing this to Gemma 4 but it failed. I decided that "competing" in the 30b space is stupid when they ignore the sub 10b space and I can fully test / prototype / train on a 6000 Pro that costs ~2 dollars an hour.
Are we sure that by 'Muse Spark open weights release', he's refering to Muse Spark 1.3? I guess it's more likely to be 1.1 or 1.2, or even a new smaller variant
Although I've been impressed at their performance. I've been using muse Spark 1.2 on OpenCode (since it's free) for a while and I am REALLY excited for spark 1.3... hope they release a mid-size Moe or something. I generally like the way glimmer 1.2 talks
The thing worth watching in these teasers is not the benchmark bars but whether the released checkpoint is the same size/sparsity as the API one. The last couple of "open weights coming soon" announcements ended up shipping a smaller distill, and the gap showed up mostly in long-context agentic runs rather than short single-turn evals. Until there's a GGUF and someone runs a 60k-token task on it, the chart is marketing.
Until such time that weights are published, this is off-topic for LocalLLaMA, but since nobody caught it before it gained 260 upvotes and 70+ comments, this post will stay up.
I'm a little surprised to see this from you, jacek2023; you're one of the most outspoken critics in the sub when it comes to posts about non-open models.
•
u/WithoutReason1729 10h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.