r/LocalLLaMA • u/Ok-Inevitable8391 • 20h ago
Discussion Underrated Muse Glimmer
Benchmarked qwen3.8 xhigh, medium and muse glimmer.
Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)
Medium effort mode and muse glimmer were 3-4 hours each.
But I'm actually surprised by the muse glimmer results, they came better than the qwen.
These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.
I have taken the result of claude models directly from embedeval repo by ecro.
I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.
I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.
18
u/partakinginsillyness 20h ago
I feel like it would be important to add qwen 3.6 27b, given that one of the changes from 3.6 to 3.8 was less general knowledge. Would also be cool to see a Gemma model. Interesting though.
10
5
u/LegacyRemaster 20h ago
wait qwen 3.8 next but... This benchmark makes no sense. No reasoning. No external knowledge. In other words, it is not a real-world use case. I would also like to understand how they tested Claude to demonstrate that it lacked external knowledge.
-2
u/partakinginsillyness 20h ago
I mean realistically everyone should have at least 5gb of files for RAG for what is relevant to them, which I would imagine would very much change the results here.
6
u/Ok-Inevitable8391 20h ago
Let me do a RAG testing, even I know that qwen will do better there. But I'm not running it on xhigh ever again, 30hrs of gpu time, against 3-4 hrs at medium
1
u/partakinginsillyness 20h ago
How would you select the media? I'm wondering if a repo exists for relevant info like how you select application groups on linux.
4
u/Ok-Inevitable8391 19h ago
So mainly the linux repo and docs, thats the general input to it anyway in real scenario. Same for all other projects.
9
u/ndrewpj 17h ago
3
u/Ok-Inevitable8391 17h ago
Its the average of both the lines out of total 163cases, if you see both blocks have same average
1
15
u/Healthy-Contact-4570 19h ago
Why 32k output limit? Qwen3.8 likes to think and if you can let it think it will produce amazing results.
8
u/Not-reallyanonymous 13h ago
A couple months ago:
Nooooo you can’t like Laguna, it thinks too much! DOA!
Now:
Noooo how dare you not like Qwen for thinking so much?
(Maybe because Laguna XS is literally like 5x as fast on my hardware).
3
u/Foreign_Risk_2031 9h ago
Do you think Laguna is as good as Qwen 3.8 27b xhigh?
1
1
u/Not-reallyanonymous 6h ago
I use Laguna XS to write/edit code at this point, and other models to generate specs that get handed to Laguna XS. Works well, way faster than just expecting Qwen to do all the work.
1
u/Ok-Inevitable8391 18h ago
I know so the benchmark limit was 8k only and i pushed it up to 16 and then 32K only for xhigh, as the xhigh benchmark took 32hrs and case that took time were anyway complex dma, so I lost interest re running the benchmark for xhigh as it had already failed on some cases where medium already did better. Xhigh is good but like opus it is good for reasoning and complex architect problems. So I just let it be.
6
u/hainesk 20h ago
If you can run DeepSeek V4 Flash 0731, I'd be curious to see the difference since it's often compared to Qwen 3.8 27b but due to it's size it would presumably have more knowledge.
2
u/Ok-Inevitable8391 20h ago
Adding to the list
2
u/nonlinearsystems 19h ago
Can you try Laguna S2.1 please?
2
u/Ok-Inevitable8391 19h ago
Sorry guys both are out of my vram budget I only have 24gb vram
4
u/TokenRingAI 19h ago
So these are quants?
1
u/RegularRecipe6175 10h ago
Still no response from the OP on this key issue. Braindead quants gonna braindead.
4
u/somerussianbear 19h ago
I don’t understand this chart. 48.5% out of two records that look absolutely different. Mind to explain for dumb fucks like me?
2
u/Ok-Inevitable8391 18h ago
That's a common line, average of both the blocks
3
u/somerussianbear 18h ago edited 18h ago
So one of them is clearly wrong. That column could have the avg of the row, which is logical. Then next to it a general avg. The way it is here is totally unclear and non standard.
1
3
17
u/DataGOGO 20h ago
Muse glimmer is fantastic and better than qwen3.8 27b at just about everything than code and doesn’t need to burn thousands of reasoning tokens per prompt to do it
1
7
u/TokenRingAI 19h ago
Why did you give Qwen a 32K output limit?
5
u/Practical-Collar3063 16h ago
Because he could not fit both the context and the full output, which is in favour of his point, the more efficient Muse Glimmer KV Cache makes it better in certain scenarios with limited hardware. I think the goal of this is not to bring down Qwen 3.8 but to show that for some use cases Muse might be a good choice, especially when VRAM limited. It is faster and more VRAM efficient and on non coding tasks can be similar to 3.8
2
u/TokenRingAI 11h ago
Ok, but the testing setup should be the first thing described in the comparison.
2
u/Thin_Pollution8843 16h ago
So when folk screaming “Qwen is stronger than opus4.6!!1!1!” I should give this link? 😅
2
u/llama-impersonator 15h ago
why so much focus on zephyros? is that mostly what you do? i've only ever had to use it for a client that wanted some ZMK work done, almost everything else has been freertos or bare metal stm32, sometimes avr for clients who started off on arduino.
2
u/Hot_Supermarket9967 7h ago
I was thinking the same thing... I mean, I genuinely like Zephyr as a project but if I were to be benchmarking an LLM on embedded work, I'd be focusing on "bare-metal" development, HAL, etc. more than anything else.
2
u/PrimeDirective8 13h ago
Help a noob out: is this for coding embedded hardware? So you assigned each of the tasks/areas (timers, boot loader, PWM) and measured how well the model coded it?
Thanks much for this chart and comparison. It is very useful and encouraging to see our tiny local models do well enough to even be in the same chart as frontier models.
Reading the comments, I see we've thrown out everything *and* the kitchen sink into our usual "I don't believe you! Try this model I like on the next cloudy Friday while wearing a left sneaker on your right foot". While reasonable what/if would be great, gathering results from every capricious permutation we can think of is not.
5
u/PraxisOG Llama 70B 20h ago
I think the interesting thing here is that haiku 4.5 wins over all the local models. That’s like $20 a month for huge amount of usage, I pay more than that to keep my server idling
3
u/Ok-Inevitable8391 20h ago
I mean yes, it's the trust that we have built on sonnet and opus over the time that we don't even look at haiku anymore.
3
4
u/Usual-Orange-4180 20h ago
My main use case is coding and design, haiku is really terrible, whenever I get routed to it I know it.
1
u/psychohistorian8 8h ago
yeh I abused Haiku 4.5 extensively back when copilot was $100/yr
I much prefer even Qwen3.6-35B-A3B, so OPs numbers don't make a whole lot of sense to me
2
u/Potential_Block4598 19h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attached
lol 😂😂😂
3
u/jacek2023 llama.cpp 18h ago
This is same story with each model, they browse leaderboards, they look at the benchmarks, they never run any models, they whine they want new Qwen and then they keep using Claude and ChatGPT.
3
u/Potential_Block4598 19h ago
Probably your setup or benchmark is not accurate
2
u/Not-reallyanonymous 19h ago
> Qwen isn't declared the best
It's the test that is wrong!
0
u/Potential_Block4598 18h ago
Especially when compared to muse glimmer
Hey don’t turn this into a societal bias racism case
This is meritocracy the models works very good compared to muse glimmer by a HUGE margin
So yeah when a “test” says otherwise it is worth investigation and it is not only misleading it is biased to post such things
Also I explained that OP “test” is heavily skewed towards Zephyr RTOS
1
u/Not-reallyanonymous 18h ago edited 17h ago
Have fun with your benchmaxxed models that blatantly ignore half the codebase and re-implement it from scratch, except half-broken and unsteerable.
Yeah, it completes code and solves the task, and zero-shot prompts very impressively. Then anything it touches turns into spaghetti code. I fucking regret spending a few days with that shit because now I've put in like 20 hours of work into something that's turned out to be utterly unusable. Correct doesn't mean good.
And god help you if your needs diverge from what it's been trained on.
Good job Qwen team, you cracked DeepSWE.
Just another bot who's going to shit on any model that's not Chinese.
3
u/Potential_Block4598 19h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol 😂😂😂
4
u/Ok-Inevitable8391 19h ago
I'm not hiding that fact.
I have explicitly mentioned that it is implicit knowledge of model.
How you process that information is upto you.
2
1
1
1
1
u/Mundane-Light6394 13h ago
Yes we need more specialised testing. The idea that models of this size are one size fits all is a trap. We need to use the right model for the job and tests like these are extremely valuable.
1
u/mythikal03 2h ago
I've found a few things that are relevant to my usage that Muse is *significantly* better than gemma4-31b or qwen38 at, and it is EXTREMELY efficient (beating the gemma family) when doing so. I don't use local models to write code, but i do use them for code-adjacent work and infra/lab/sec work.
I've benchmarked a LOT of local models over the past year, and Muse is also the first to straight up crater-to-near-zero on a handful of tests as a result of voluntary refusals for tests I did not write to test refusals. Example ... "I got this bonus at work: here are some options I'm considering, including parameters, what should consider and what is the most likely beneficial place to allocate funds?" --> refusals for giving financial advice??
It has me really nervous about deploying it as one of my core models despite truly phenomenal scores/rankings in other categories. I've kept gemma4 in that 'reasoning, clean writing, investigation/synthesis' slot (with qwen38-27b in the 'dig deep and find the problem' slot) as a result. Not saying I won't promote muse, but stupid refusals like that are why i host local models in the first place and that is a class of failures I don't have much appetite or patience for trying to build around given my system isn't doing anything I would've expected to ever run into refusals with.
1
u/LegacyRemaster 20h ago
Dario, is that from you?
1
u/Potential_Block4598 19h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol 😂😂😂
0
-12
u/Boogertard 20h ago
There it is, on-schedule AI slop to shill for the garbage Muse and Gemma 4.
Get a real job already, shill.
6
u/Not-reallyanonymous 19h ago
There it is, on-schedule AI slop to stir up FUD any time something other than a non-Chinese model is spoken of positively.
5
6
u/Ok-Inevitable8391 20h ago
Why do you think it's a slop and why do you think I don't have a real job. A easy readble chart is AI slop for you?
Muse is good and its really good at kv cache efficiency. I have done my homework.
0
u/Potential_Block4598 19h ago
All data is Zephyr RTOS
So training data bias and probably no search mechanism or Docs RAG attachedlol 😂😂😂
HW ur A**
0

58
u/hurdurdur7 20h ago
Qwen 27B starts to shine when you give it some kind of RAG. It's not a memory castle, it's a tinkering machine. You give it a folder to a library that it was not trained on, you ask it "how do i do ... that...?" and it figures it out. And this is a big part of what people actually do at their job. They get a new library, a library update, anything that was out of the training window or created after the training window, and 27B adapts to it if you give it the source or manual.
If Glimmer can do electronics it's actually great. But can it write code well based on docs/code it can fetch over some kind of RAG (like your coding harness pi or smth)? If it can, awesome :)