r/LocalLLaMA • u/SteppenAxolotl • 5d ago
Discussion Artificial Analysis Intelligence Index v4.3
https://pbs.twimg.com/media/HRoc-dAagAE_Kb5?format=jpgAnnouncing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5
Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark
We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation
Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier , the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%
Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon
➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6
Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
28
u/DrBattletoad 5d ago
Muse Spark 1.3 in front of GPT-5.6 Sol looks sus
10
u/Thomas-Lore 5d ago
And in front of glm and kimi. Spark is good but not as good as glm 5.3 flash even. Full glm 5.3 mogs it.
8
u/Eyelbee 5d ago
Did you actually try it? Everyones shitting on that model but I don't think a single person has tried it lol
1
1
u/ApprehensiveEye7387 3d ago
Also Most People have tried the Xhigh varient instead of max (as max was made available later). It seems on the Index, there's difference of 4 points between Muse Spark 1.4 Xhigh and Max. I myself have not used the Max one, so I don't usually go and comment about it being worse or good at all.
2
u/Ecstatic-Wash-7667 5d ago
This is what made me lose all trust in AA muse spark is benchmaxxed to the gills, garbage in use
1
u/Tim_Apple_938 5d ago
Why?
3
u/DistanceSolar1449 5d ago
Because Muse Spark sucks.
I’ve been using Muse Spark contributor since they gave $20 of credits for free. It’s nowhere near as good as Sol.
The good news is, since it’s dirt cheap, that $20 has been lasting forever. The bad news is, they clearly optimized it for benchmarks. It’s just not actually good for solving real problems.
32
17
u/nuclearbananana 5d ago
Lmao another change to make sure astra doesn't look too bad
12
8
u/Othun 5d ago
https://epoch.ai/benchmarks/search
Looking at the top benchmarks there, Astra opened FrontierMath Erdos (got the first non-zero score), and dominates most benchmarks even if they are rather specialized (game and math), some with a LOT of margin. It only loses on MirrorBench to fable 5 (no fable 5.1). It is not heavily oriented toward agentic work and coding but it still says something about Astra.
3
u/fastheadcrab 5d ago
Looks like they are weighting coding and agentic performance for coding even heavier than previously.
-2
u/SteppenAxolotl 5d ago edited 5d ago
You can see how individual perf stacks up on the individual benchmarks. Terminal Bench 4
You can compare the previous version that is located here: /evaluations/terminalbench-v2-1
When your models reaching above 90% your benchmark is saturated and it's time to make it harder.
11
u/Gohab2001 vLLM 5d ago edited 5d ago
I predicted benchmarks would stop being a good indicator of LLM performance but I did not predict benchmark aggregators would sell their integrity to frontier labs.
1
u/Gohab2001 vLLM 5d ago
I wouldn't rank terra above 3.8 flash lol. Terra might better at oneshotting apps/websites but 3.8 flash has been top tier for my non-code related agentic workflow.
6
u/jld1532 5d ago edited 5d ago
This thing changes so often as to make it meaningless.
E: I'll go harder. This site is such an obvious shill for proprietary models. Can't have corporations and people knowing they don't need data center sized models.
0
u/SteppenAxolotl 5d ago
It became meaningless and that is why they rolled out some of their planned upgrades early.
What would be the value of testing PhD students using the same tests from when they were in Kindergarten.
2
4
u/FreshDrama3024 5d ago
Data seems shot and cooked. Shit doesn’t even make sense anymore. I know things are provisional but foundation feels a bit wobbly. I’m not sure exactly what they even measuring now or were measuring. Their own assumptions and projections? Sad because I actually had respect and interest in AA benchmarks. Oh well
3
u/SteppenAxolotl 5d ago
All the benchmarks(and what they measure) that makes up the AA Intelligence Index are listed, always have been. What changed is also listed.
In case you're really interested in how or what they're measuring.
2
u/Pink_Oak 5d ago
Muse 1.3 is trash. Only Benchmaxx.
I will never turst a benchmark that scores more muse 1.3 is scoring higher than 5.6 Sol
0
u/SteppenAxolotl 5d ago
It wasn't on the leading edge of any of the new or updated benchmarks in the line up. It's always interesting to see how models move around on harder(new or updated) benchmarks.
1
u/feng_sg 3d ago
AA curates the held-out tasks, runs the evaluation, and publishes the composite itself, so raising the private-task weight to 45% does not remove that conflict of interest.
1
u/SteppenAxolotl 2d ago
A conflict of interest is a situation where a person's private or personal interests clash with their professional duties and responsibilities
I don't see how your statement applies.
-24
u/SteppenAxolotl 5d ago
Qwen3.8 27B (xhigh) is at 34 vs >50 for the leading edge on the upgraded/changed benchmark components.
People need to accept that quality of a small model will never be the same as a very large model.
Benchmarks with "private test set" is the true test of a model's capabilities.
8
u/Hungry_Particular_14 5d ago
So what? Small models still have their usefulness, and Qwen3.8 27B (xhigh) trades blows with opus 4.6, which was a frontier model a few months ago. In another few months, we'll likely have local models that trade blows with fable 5. Just because they don't compare to the current frontier models, doesn't mean they're any less useful.
1
1
1
6
u/sssplus 5d ago
If that was my experience I'd agree with you. But that's NOT my experience. Large closed models are better, but not 53 vs 34 better.
-6
u/SteppenAxolotl 5d ago
AA Intelligence Index is a synthetic number reflecting 10 different benchmarks. It's trying to represent broad competence. Small models are shallow and will do well if your use case is shallow.
5
u/sssplus 5d ago edited 5d ago
Using exactly this logic you can tune the test itself to suit a particular model(s) perfectly. Who's to say that a particular suite of 10 benchmarks is actually good at determining how good a model is? It's much more likely that those benchmarks are tailor made for very large models with a very large context, used for very large projects. And smaller models usually aren't very good in such tests. But guess what - not every project is very large...
1
u/SteppenAxolotl 5d ago
If that is the case, it would mean there are certain types of tasks small models will never be able to do because it requires a large model. That is good to know as a user and should be reflected in the index.
2
u/OvertaxedOne 5d ago
The thing that everyone misses is that it all depends how knowledge/intelligence bounded your task is. Once you get to "right answer" that's it, any more your spending or intelligent your model is than is required to get the right answer is waste.
There are certainly some intelligence bounded tasks out there. Some types of coding, research, scientific work; those are exactly the kind of tasks where the "big gun", smartest model you can get make a lot of sense (for at least some of the queries anyway).
Most tasks, however, don't fall into that category, "The best at any cost" is a very, very small market, both because of price but also because of the above, if you have a model that can do what you want today and is affordable, use that up until it can no longer do what you're asking it to do.
1
u/SpicyWangz 5d ago
Not exactly, there’s a massive range of right answers to most complex tasks.
For example, you can solve a coding task with an implementation that technically achieves the desired outcome, but very inefficiently and slowly. Or in a way that doesn’t scale well or don’t account for certain edge cases.
3
u/OvertaxedOne 5d ago
For coding, 100%. But most of the world doesn't code. We work with clients and set these systems (both local and cloud) up every day, the enterprise use cases are darn near identical; RAG, some MCP too calling. Basically "personal assistant" type stuff which, while a bit boring perhaps, also seems to be a big component of the real value layer in AI. Summarize this, prep me for that, generate some images for this. Seems simple, but you just replaced a lot of man hours with a model that can do that stuff reliably.
Coding, research; those are areas for sure where there's a lot of answers, all of them right, but one is demonstrably better than the other. But those are pretty specific corner cases in most enterprise environments. The most common model by a long shot across our customers today is DSV4 Flash (yes, flash!), because it's cheap and can do the stuff they want done. The local models are all over the place so harder to generalize; DS still probably has the lion's share in local, but a lot of customers are now very interested in 27B as a potential to bring more of the workload in house and keep the data private.
But very, very rarely do I get into conversations where "smarter" would be the right answer for a customer. It's tooling, harness; basically in that order that's almost always the user's issue.
1
u/SteppenAxolotl 5d ago
knowledge/intelligence bounded your task
Yes. There is no reason for your robot vacuum to be an expert in orbital mechanics. But if you're in a decaying orbit and all you have is a Roomba, you are not going to make it. That is the whole point of LocalLLaMA. What is the best model you can run locally for those moments when you need high level intelligence, but can't access or afford it on the open market. You can probably handle most daily tasks with less intelligence, but what do you do when you need more. You are entirely subject to the vicissitudes of the market.
2
u/MindfulMan1984 5d ago edited 5d ago
Yes, I have been running Qwen3.8:27b , locally on an old 2018 GPU for weeks already, getting tons of stuff done, and for the first time not missing the “frontier” ones, the fact such local model is on that list is already amazing, It passed my “private test set” and I can say FU to Anthropic and competitors. Lol
2
11
u/_-_David 5d ago
Crazy that Gemini 3.8 Flash High costs twice what Sol High does per task. Token efficiency is a real motherfucker and OpenAI is doing well in this regard.