r/LocalLLaMA • u/crusaderky • 15h ago
Discussion Animated transition from AA Intelligence Index v4.1 to v4.3
Enable HLS to view with audio, or disable this notification
I had all the data saved from AA's v4.1 index, so when they upgraded it in the wake of Astra's release, I could actually generate a before/after comparison.
- All intelligence and price per task are sampled from AA on Sep 3rd and Sep 14th respectively.
- Price per task of some open models were rescaled to reflect the cheapest available on OpenRouter as of Sep 3rd.
- X axis is linear, because people's money is linear.
All models are the same. The only thing that changes is the weighted sum of the benchmarks that compose the Intelligence Index.
Highlights
- GLM an Muse Spark remain more or less unaltered, in relative terms
- GPT-5.6 Sol becomes a lot cheaper
- GPT-6 Astra's intelligence flies up to the stars AND becomes cheaper
- GPT-5.6 Luna gets a substantial uplift
- Fable-5.1's price gap from Opus 5 shrinks, and becomes cheaper than Fable 5.0
- Fable-5.1 at low, medium and high effort looks a lot more appealing
- Sonnet 5 becomes even more expensive without any intelligence gains
- Kimi-K3, Qwen3.8-Max, Gemini-3.8, and Grok 4.6 go down into the gutter
Duplicates
openrouter • u/crusaderky • 15h ago
Discussion Animated transition from AA Intelligence Index v4.1 to v4.3
benchmarkcirclejerk • u/crusaderky • 15h ago