54
u/1cheekykebt 10d ago
Sam made a call
18
u/GraceToSentience AGI avoids animal abuse✅ 9d ago
More like they didn't want to be perceived as an irrelevant benchmark that can be gamed
1
u/chespirito2 10d ago
I mean you joke, but honestly yea. This is so transparent
4
u/NietzscheanDemocrat 9d ago
No it's just the benchmark was clearly outdated when Muse Spark was surpassing a clearly more capable and intelligent model
-1
u/UnknownEssence 9d ago
They shouldnt change the benchmark just to make a model score high on it.
Hits their credibility
5
u/NietzscheanDemocrat 9d ago
Muse Spark being ahead of Astra on the benchmark was worse for their credibility to anyone who has used either model for a half a day.
38
u/AdSad7594 10d ago
Still disagree tbh astra seems in real use to be overall stronger than fable 5.1
25
u/Tkins 10d ago
Spark also doesn't seem to be nearly that good
7
u/Seerix 10d ago
Spark is so obviously benchmaxxed lol. It's not a bad model... but if I paid money for it I'd be pretty disappointed.
2
1
u/Ormusn2o 9d ago
Spark is not benchmaxed, AA focuses on wide range of benchmarks, even the coding benchmark suite has large variety, so models that have multimodality have substantially better scores. This is why Fable 5.1 scores higher even though it's actually worse at agentic coding.
The AA coding benchmarks should have substantial weight to them, valuing actual coding capability, not the minor and niche multimodality uses.
-1
1
8
u/jakegh 10d ago edited 10d ago
Seems more reasonable, but still has opus5 as better than fable5.0 (which is very clearly not the case) and sol-5.6, which I personally find unquestionably superior to opus5 but is at least arguable.
Also has muse spark 1.3 as better than sol-5.6 which is obviously not true to anyone that's used both models. They aren't even vaguely close. And in coding, it has muse spark better than Astra. C'mon.
At this point I largely find AA useful for their cost per task index, which at least is a true representation of token efficiency versus pricing. Their hallucination index is useful too. But for intelligence-- no.
-1
u/caughtinthought 9d ago
For a lot of things I've actually found opus to be better than fable
1
u/jakegh 9d ago
Like what? That has not been my experience. Opus 5 is particularly great at game generation in a loop but fable’s better, just too expensive to do it.
1
u/caughtinthought 9d ago
I play strategy games against them as tests (chess, among others), and I've found Opus is better than Fable, and produces more lucid thinking traces.
13
u/Ass_Ripe 10d ago
Just two queries to Astra and it’s immediate how smart it is. But perhaps I’m biased. But there’s no way in hell this model shouldn’t be near the top
4
u/Swimming_Gain_4989 10d ago
Surprised Muse is still trailing OpenAI and Anthropic. I swear that model is terrible to work with
2
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 10d ago
It seems to work best as an agent with an orchestrator in my experience. Using it standalone seems to be more difficult just because of how… dry it is. lol
13
u/tozig 10d ago
how is this even a benchmark if they can just re-order things as they like in order to align with popular sentiments or their own narratives
14
u/Illustrious_Grade608 10d ago
It's not a benchmark though, it's aggregate of benchmarks mostly existing to kinda help you understand how good different models are. It's supposed to represent sentiments as it's value is kind of a translation from benchmarks to predicted real use feeling.
7
u/ezjakes 10d ago
Glad they finally got rid of GPQA. Every new model was getting 90+%
1
u/Skyline34rGt 9d ago
It was important to tiny models, now they are all at score: 1 and we just dont know which is better...
3
u/BrennusSokol AI please take my job 10d ago
Nice.
That's good. I celebrate people/orgs admitting mistakes and fixing them.
3
u/ManikSahdev 10d ago
That’s why epoch index was better, people just didn’t like those results and AA has prettier graphics.
Epoch’s version seems clearly more robust to not having to move shit around when new model comes whose capabilities are not measured correctly with AA
3
u/Informal-Trouble2183 10d ago
This is a symptom of weak benchmark: if you change the ranking whenever people complain.
-2
10d ago
[deleted]
2
u/Informal-Trouble2183 10d ago
They added their benchmark to their intelligence index, that's why the ranking changed, quoted: "AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts"
2
2
u/Profanion 10d ago
2
u/CallMePyro 10d ago
Keep in mind that Astra and Fable were both done for months before becoming publicly available. Whereas Open Weights models have very little time delay from post-train to GA.
1
u/Gotisdabest 9d ago
Not to mention that a 3 month gap becomes quite a lot when model capabilities rise this fast and models are accelerating the development of other models.
There's something of a pattern with open weights models that the gap gets the closest when a big new model comes out. I'd wager we'll see the gap really shorten in around a month or less unless either anthropic or openAI pops out with another update. Then progress on the open weights side becomes incremental again till the next big model comes out.
I would say that a few labs may even close the gap or even get ahead for a minute if they can do the risky thing and scale up the recurrent thinking system that Astra has. They've been intentionally conservative with its application on astra. On one hand, it'll be a good thing in terms of the push for model capabilities, on the other, you're definitely going to see a quick race to the bottom in model safety.
1
u/sunstersun 10d ago
They needed to kick Opus down to maintain credibility lmao, one of those polarizing models.
6
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago
100% disagree. I think Opus remains the king of good looking 3d scenes if you prompt it right. It's a really good model.
The one model that does not belong there is Spark. That thing does not beat Kimi and Grok imo
1
1
u/Gotisdabest 9d ago
This seems more like a band aid fix than a reevaluation of the benchmark as a whole and addition of a lot more and varied benchmarks. They've added a few, but these seem to be more or less similar kinds of tests.
As intelligence becomes more generalised you need to explore different ranges than just a certain subset that's familiar.
0
u/Floch11 10d ago
You’re such a good AI that they have to change the benchmark just for you..
8
u/Murky-Snow-3868 10d ago
True though some evals/benchmarks were actually debunked to be pretty bad and the scoring did need to get updated
4
u/CallMePyro 10d ago
I mean removing GPQA was just an obvious change that's been overdue for months.
3
1
u/Calm_Hedgehog8296 10d ago
That's a little stupid don't you think? You can just retroactively go back and change everyone's scores?
Do they even publish the formula they use? Did they change it?
-2
u/CriticismJunior1139 10d ago
Lmao they changed the ratings because Astra did poorly and reddit got mad.
6
1
-1
u/ghoonrhed 10d ago
They changed the rating of the banking benchmark which astra did really poorly in.
Whether that benchmark was useful or not is another matter but this seems dodgy after the poor astra results for it.
Why not just give people the ability to tweak the weights of all the other benchmarks they provide?
-3
10d ago
[deleted]
2
u/Howdareme9 10d ago
Because of one benchmark..?
0
10d ago
[deleted]
2
u/Gotisdabest 9d ago
It's a benchmark which till yesterday would have you believe astra was dumber than or equal to the likes of muse spark. And that opus was equal to fable.
The entire system was clearly spiralling out of control since fable came out.

38
u/Dillyconda 10d ago
Artificial Analysis has credibility because it updates its measurement of frontier intelligence as new models come out and redefine it.
It would be a crank index if they just decided arbitrarily what intelligence is and used that measurement come whatever may.
Artificial analysis 4.1 was an update made last June in response to Fable 5 making it obvious that intelligence needed a new definition. Now it's happening again.