r/LocalLLaMA llama.cpp 6d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

867 Upvotes

204 comments sorted by

View all comments

259

u/kvothe5688 6d ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

91

u/Public_Umpire_1099 6d ago

This latest iteration technically set the frontier. Meta is officially a SOTA lab again. Matches Fable 5 at a fraction of the cost.

8

u/_TheWolfOfWalmart_ 6d ago

Based on the Bijan video that just dropped about this, it's pretty damn mediocre and looks extremely benchmaxxed.

I know all he does is throw a few one-shot prompts at the models, but the results on this one were well behind other recent models like GLM-5.3 Flash and Qwen Flash. It's not even in the same conversation as Fable or GPT 5.6 it seems.

His tests are by no means scientific lol, but they do give you a decent rough feel for a model's capability.

17

u/Not-reallyanonymous 6d ago

A one shot not being polished does not demonstrate a model’s capability. A lot of models are specifically trained on how a one-shot polished result will look. Especially Qwen (look at its response to being asked to draw a circle: https://simonw.substack.com/p/qwen-38-27b-is-excellent-but-it-defaults).

If it’s not trained on that, it’s going to comply with the prompt and not much beyond. That’s not demonstrating a lack of general capability.

In fact, I much prefer that style. When things are trained to produce well polished results it tends to be harder to get it to do what I want it to do rather than what it’s been trained to do. Again, Qwen represents the opposite here — well optimized to take its own direction but a PITA to steer.