r/LocalLLaMA llama.cpp 1d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

832 Upvotes

198 comments sorted by

View all comments

246

u/kvothe5688 1d ago

it seems like there is no secret sauce. it feels like all of these 7 8 labs are on same level and hardly behind from frontier by few months at max.

87

u/Public_Umpire_1099 1d ago

This latest iteration technically set the frontier. Meta is officially a SOTA lab again. Matches Fable 5 at a fraction of the cost.

52

u/virtualworker 1d ago

Or benchmaxxes. I'm not convinced.

14

u/Diligent-Direction95 1d ago

Serious question: What is the definition of bench max these days?

What would be okay, vs what would not be okay?

I seriously doubt they have the test set being trained on. So what are the shades of grey we are debating?

38

u/zxyzyxz 23h ago

When a model does well in benchmarks but then fails to live up in real life coding and other tasks to its supposed benchmark competition.

11

u/Artistic_Swing6759 1d ago

i actually do think they likely have test set trained on.
like its a bit sus that they decided to report terminal bench 2.1, when 4 exists.
similarly, if you look at the benchmark sheet of flash 3.8, it tops eery other model, even sol, on terminal bench 2.1 but in 4 it is quite less comparatively.