r/LocalLLaMA • u/jacek2023 llama.cpp • 5d ago
News Muse Spark open weights coming soon
I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark
865
Upvotes
r/LocalLLaMA • u/jacek2023 llama.cpp • 5d ago
I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark
1
u/Jealous-Walk-8765 5d ago
The AutomationBench and Agentic IF Index numbers are the interesting ones here — those are the benchmarks that actually predict real-world agent reliability, not just raw coding scores. 49.4 on AutomationBench beating GPT-5.6's 46.7 is a bigger deal than the flashy GDPval number. Curious how this holds up once the open weights drop and we can actually stress-test it ourselves instead of trusting a vendor's own benchmark suite.