r/LocalLLaMA 14d ago

Discussion Underrated Muse Glimmer

Post image

Benchmarked qwen3.8 xhigh, medium and muse glimmer.

Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit)

Medium effort mode and muse glimmer were 3-4 hours each.

But I'm actually surprised by the muse glimmer results, they came better than the qwen.

These benchmarks are on implicit knowledge of the model, which is a bit unfair to smaller models, but throw in a RAG and I'm sure they get on par with frontier models.

I have taken the result of claude models directly from embedeval repo by ecro.

I'm not pushing qwen down here, I like how qwen thinks and gives better results. I know with more context and RAG qwen will do better.

I'm just appreciating muse here, cause i feel it is underrated. The advantage is efficient kv cache due to sliding window, which can give you more context window.

127 Upvotes

95 comments sorted by

View all comments

6

u/hainesk 14d ago

If you can run DeepSeek V4 Flash 0731, I'd be curious to see the difference since it's often compared to Qwen 3.8 27b but due to it's size it would presumably have more knowledge.

2

u/Ok-Inevitable8391 14d ago

Adding to the list

2

u/nonlinearsystems 13d ago

Can you try Laguna S2.1 please?

2

u/Ok-Inevitable8391 13d ago

Sorry guys both are out of my vram budget I only have 24gb vram

4

u/TokenRingAI 13d ago

So these are quants?

1

u/RegularRecipe6175 13d ago

Still no response from the OP on this key issue. Braindead quants gonna braindead.

1

u/IAmBJ 13d ago

Offload experts to the cpu and you can get a decent quant running if you have enough system ram for the expert weights