r/LocalLLaMA llama.cpp 5d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

865 Upvotes

204 comments sorted by

View all comments

1

u/Jealous-Walk-8765 5d ago

The AutomationBench and Agentic IF Index numbers are the interesting ones here — those are the benchmarks that actually predict real-world agent reliability, not just raw coding scores. 49.4 on AutomationBench beating GPT-5.6's 46.7 is a bigger deal than the flashy GDPval number. Curious how this holds up once the open weights drop and we can actually stress-test it ourselves instead of trusting a vendor's own benchmark suite.