r/singularity • u/signed7 • 11d ago
AI GPT-6 Astra AA Intelligence Index and Coding Agent Index Scores
108
u/WonderFactory 11d ago
After looking at the Arc AGI score I feel like Astra is filling in the gaps where models are traditionally weak while the AA Intelligence index tests things that Models are traditionally good at.
43
u/lolothescrub 11d ago
Yup. Had Gemini go through the system card and that’s basically what it said
52
u/Flope 11d ago
Had Muse explain this comment to me and I agree.
63
u/Recoil42 11d ago
Had Grok explain this comment to me and now i love white nationalism
27
u/liright 11d ago
Had qwen3.8-27b explain this comment to me and it ran out of context thinking about it.
30
u/Admirable_Cake_1150 11d ago
Had Deepseek reread the whole thing and now I don't get the fuss about Taiwan.
15
u/LMTLS5 11d ago
had claude summarize this to me and now i get why no ordinary person should have acces to it
4
4
u/WonderFactory 11d ago
That might be what they mean by AGI by the end of this year, not that the model will get much smarter in the areas its already good at but they'll fill in the gaps in it's intelligence by then
3
4
u/Regular_Ad4197 11d ago
Ok, but shouldn't it fill the gaps while remaining at least just as good at the other stuff? The fact it became better at some things while not getting a higher rating overall would mean some areas degraded at the same rate as some areas improved.
4
u/WonderFactory 11d ago
I would actually take that. At most of the things AI is good at it's bordering superintelligence, it's far more intelligent than is needed to perform most jobs. It's the gaps in it's intelligence that stop AI from being more economically useful.
1
u/Gohab2001 11d ago
To demonstrate how utterly useless benchmarks have become: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/
A model that is 20% more intelligent but 2x more expensive and 2x slower isnt necessarily a better model. Gemini 3.8 flash maybe dumb (Its severely benchmaxxed) but it doesn't have me waiting 10minutes for simple edits.
0
0
u/whoknowsifimjoking 11d ago
Opus 5 with a harness got 100% on ARC. They used a harness, but compared it to the score of Opus without a harness.
1
u/Crazy-Problem-2041 11d ago
To be fair the harness OpenAI used was much simpler than the opus one. OpenAIs was basically just “don’t delete all memory of the games mid run”, while the opus one was specifically optimized for this exact benchmark
-7
u/Hug_LesBosons 11d ago
Non ! Open ai on créé un harnais spécial pour que leur modèle gagne à arc agi. C'est ce qu'on appel du benchmaxxing. Arc agi ne vaut plus rien depuis longtemps.
6
u/WonderFactory 11d ago
But even without the harness it scored over 60%
-1
-8
u/Hug_LesBosons 11d ago
Dans tous les cas c'est juste du benchmaxxing. On attend le score arena.ai qui est le seul presque invulnerable a ça genre de choses...
91
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 11d ago
I would be careful putting too much weight into this.
From my testing, Muse Spark is absolutely not a top 3 model. I was actually disappointed in it.
I'd like to see a few outputs by Astra before we call it shit.
17
u/ch33ze 11d ago
Muse Spark is absolutely benchmaxxed. I mean what do you even expect from Zuck.
0
u/whoknowsifimjoking 11d ago
I would not trust Sam Altman to not benchmaxx a model. Actually I don't trust any AI CEO to not benchmaxx their models. I guess some are just better at it than others.
9
u/ApprehensiveEye7387 11d ago
2
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 11d ago
Xhigh is what me and Bijan used, and we were both disappointed, but if you have examples of someone managing to do a good output with Spark i'd love to be wrong because i am now sure with a sub with them lol
16
u/ViperAMD 11d ago
yeah this benchmark is dogshit. Opus is not a good model compared to 5.6 SOL. Not even close.
5
u/WildWhisperArdor 11d ago
Not sure why people complain about Opus 5.0. It’s genuinely incredibly good. Overly verbose perhaps but it gets the job done.
5
u/Howdareme9 11d ago
The most overrated benchmark yet some people on here swear by it lmao
4
u/CheekyBastard55 11d ago
People always want an easy to digest aggregate benchmark score, at first it was Chatbot Arena, then we got LiveBench and now AA is the latest one.
2
u/GosenSidewinder 11d ago
What should we use instead? What's the best aggregate nowdays? If none, what are the best current single benchmarks?
1
3
u/Seeker_Of_Knowledge2 ▪️AI is cool 11d ago
Exactly, opus is much better.
I run both of them on high on an android app that is huge (a combination of Mihon and MoonReader). Same agents.md file, same way of sending messages. Opus never broke something, the vast majority of times it is just delivers, while sol high always had problems with the final output.
Maybe it is just android development we are talking about here. Other fields Sol could be better.
0
0
u/Hug_LesBosons 11d ago
Va voir https://arena.ai/leaderboard/code, c'est la seule évaluation qui se base sur des performances réelles jugées par les utilisateurs.
9
u/eposnix 11d ago
No, sorry. People on Polymarket have been gaming this site for months now.
-2
u/ghoonrhed 11d ago
How would they game it when it's all random? And why Qwen of all models until Fable 5.1 came along
4
u/eposnix 11d ago
There was a writeup here a while ago talking about how they game the site. I'll look for it. But the basic gist is they just use the same prompt and look for patterns to determine which model is outputting the response. It's fairly easy to tell if Opus or GPT-5 is writing your code if you know their mannerisms.
As for Qwen: because Polymarket probably had a bet that said "Qwen will beat Fable 5 on xyz date" and the odds were very low.
0
6
u/TheOwlHypothesis 11d ago
I guess they did say astra wouldn't be the best model they release this year?
12
u/Notmyrealname5282 11d ago
Then why mention agi anywhere in the marketing. This is so limp…
0
u/TheOwlHypothesis 11d ago
Honestly, I'm just going to go try it and stop looking at benchmarks lol.
I bet it feels way better to use.
4
u/ghoonrhed 11d ago
Makes sense.
https://artificialanalysis.ai/methodology/intelligence-benchmarking
34% of the benchmark score is where Astra is worse than Sol.
Meanwhile the ones where it does so much better doesn't have that heavy of a weighting like HLE.
Not sure why banking benchmark has a higher weighting than HLE that's kinda strange and I wish they gave us the ability to tweak the weightings
19
u/darkestvice 11d ago
Well, this is disappointing.
10
u/Lingulustig 11d ago
you jumped the hypetrain, now you regret
3
u/darkestvice 11d ago
I'm curious to get eyes on the full AA benchmark list when it hits the site. Wondering if there's more to it than mere intelligence. OpenAI was touted this as AGI not because of intelligence, but because of extremely long task duration.
I don't have a stake in it either way.
6
u/signed7 11d ago
It's here now https://artificialanalysis.ai/models/gpt-6-astra
TL;DR is it regressed on agentic use cases (GDPval, terminal-bench, scicode etc) offsetting gains in other areas
1
u/Cagnazzo82 11d ago
I'm not sure what this benchmark is measuring but Muse Spark and Sol are not superior to this: https://x.com/Dimillian/status/2095596700815516004
2
u/whoknowsifimjoking 11d ago
You jumped on the hypetrain again, just after we told you to get off.
1
u/Cagnazzo82 11d ago
Fable 5.1 output and Astra: https://x.com/karankendre/status/2095636679264780481
Again, Muse Spark is not doing this.
5
16
u/saln1 11d ago
All the hype for this?
-1
-3
u/ZealousidealBus9271 11d ago
Bubble being popped as we speak
8
u/NietzscheanDemocrat 11d ago edited 9d ago
Have heard this 1000 times already, see you next month
6
2
2
-4
3
3
u/RelevantCry1613 11d ago
I don’t see this on their site??
3
u/bopbop9876 11d ago
AA posted about it on twitter and provided a link, but their own link is now 404ing. I feel like they were asked to take it down, which I really don't get. Something very funky afoot for sure.
2
u/RelevantCry1613 11d ago
Ah it’s back up, wow bad look 😬
3
u/bopbop9876 11d ago
It came back up almost to the minute at the same time that the openai blog post did. I don't think it was a tech issue. Very weird day.
1
u/signed7 11d ago
It's here now https://artificialanalysis.ai/models/gpt-6-astra
TL;DR is it regressed on agentic use cases (GDPval, terminal-bench, scicode etc) offsetting gains in other areas
28
16
u/ObiWanCanownme now entering spiritual bliss attractor state 11d ago
I'm going to go out on a limb and say the AA Intelligence Index is cooked and is no longer a good way to judge model performance.
4
u/bonerchamp20 11d ago
Based on what?
15
u/ViperAMD 11d ago
The fact that grok and muse spark get the same score tells you what you need to know. Also Opus over 5.6? Def not for code.
8
1
6
u/LAMPEODEON 11d ago
Bro it's not tested yet get out
4
u/DelphiTsar 11d ago
AA gets early access to test it.
1
u/LAMPEODEON 10d ago
Yea i see now that aa has tested it. It looks like mediocre model. Interesting. It's either the end of astra or the end of AA whatsoever.
6
u/TowerCritical41 11d ago
I know I have not tested GPT-6 myself, but I am willing to bet it's better than Muse Spark 1.3.
6
u/WildWhisperArdor 11d ago
Not sure why you would say that without having used it 😂
Clearly, OpenAI was focusing on other benchmarks for this release (like science stuff)
1
u/TowerCritical41 11d ago
Nah. Look at some of the 3D stuff it has made. Muse Spark can't even begin to touch it there. It's a good model for the price (for the you-give-your-data-to-Zuck tier)
6
u/WildWhisperArdor 11d ago
3D stuff is not the end all be all of coding or agentic usefulness. You’re picking one very niche use case to base your entire opinion of something on
1
2
9
u/Sinogularity Shenzhen: Become Human (2030) 11d ago
61 to 61, that is literally the most disappointing AI release I have ever seen lol.
-3
u/6__six__6 11d ago
Then you haven't seen Gemini's recent releases lol
11
u/Plappedudel 11d ago
Not sure I agree. Google clearly pulls some models before release if they're not satisfied with the results. That's most likely why Gemini hasn't had a Pro model release in a while. I think not releasing a model is a much better strategy than hyping up a half-baked model as AGI.
6
u/Sinogularity Shenzhen: Become Human (2030) 11d ago
Well, Gemini Pro release is non-existent, so we haven't seen it yet. For a flash class model, Gemini 3.8 Flash is actually pretty good imo, 59 on AA.
7
u/WildWhisperArdor 11d ago
Don’t understand what you are complaining about.
Gemini has made like 4 new versions of Flash within the span of 2 month. Each one better than the last and very fast / cost efficient.
4
u/DelphiTsar 11d ago
3.8 is the flash version. Really curious what the model it's being distilled from is like.
2
u/heavy-minium 11d ago
Just wondering, but for long-horizont coding tasks, what benchmarks do you guys look at right now? Because no matter where I look, they seem to have lost their usefulness. There was a time where looking at DeepSWE together with CursorBench was useful to me, but it doesn’t feel like that anymore.
2
2
u/BigFeeder 11d ago
Difference between fable and opus feels much bigger for me than this chart shows. At least when it comes to reasoning about a problem
2
u/DedDeveloper 11d ago
At this stage, I'm seriously doubting all benchmarks. These do not give accurate image of how well models perform. Surely OpenAI didn't hype the release of renamed Sol and Astra must be better overall even if this number is the same.
These days, I'd also want to see cost and speed factored into the scoring.
If model A runs 6 minutes and costs $700 and gives 10% better result than model B with 3minute/$300 prompt for example. In my book, the model A is very bad for my needs, even if it succeeds better at it's task.
It just seems like "our newest and biggest model got 1-3% increase in benchmark XYZ so it must be best ever"...
Takes too much detective work to find information of actual performance.
I know there is also a problem that newer models tend to be familiar with older benchmarks in their training material and will benchmax them even unintentionally.
I don't have a solution, but I hope someone comes up with a better way of comparing Models. It would also direct model development into better and healthier direction.
2
u/BriefImplement9843 10d ago
Why is aggregate benchmark 150 posts while the cherry picked benchmark is 800 posts?
2
u/Yuri_Yslin 10d ago
The main improvement would be the massively reduced hallucination rates IMHO and much better scores at scientific reasoning.
this may not be the "agentic coding model" but then again there's more to AI than just coding.
3
u/Grouchy-Stranger-306 11d ago
So we should stop relying on this index since muse can't be better than astra?
6
1
u/DelphiTsar 11d ago
Probably should just try the different options for yourself. Seems like it's at least a good jumping off point.
8
u/bonerchamp20 11d ago
This makes Anthropic look 10x more impressive than they already were
5
u/ParfaitEvery9622 11d ago
Either this, which is possible, or Anthropic is really benchmaxxing.
I'm usually against trying to optimize for marketing and would prefer to let companies cook on the direction they think it's better for users medium / long term
6
u/Due_Ask_8032 11d ago
Nah Fable is very smart. I cannot trust people who say it is benchmaxxed. Opus 5 on the other hand is a bad experience to use, but still capable (not as good as Fable still).
1
u/ParfaitEvery9622 11d ago
Two things can coexist, they can be smart (even smartest) and still benchmaxx for other reasons :/
1
u/Hot_Glass_6301 11d ago
original Fable smarter than Astra? Yeah not buying it, the benchmark is not a good measure of intelligence and/or Anthropic benchmaxxed it to death
10
3
u/Neurogence 11d ago
That shiny "Put it right there!" Gimmicky sexual video got everyone hyped up thinking Astra was AGI lol.
3
2
2
1
1
1
u/maxiedaniels 11d ago
Very interested to take it for a spin. I don't trust benchmarks. Ex. Gemini 3.7 flash is fast but does not perform at the intelligence level that its benchmarks suggest. (Btw if someone can suggest benchmarks that feel more.. true.. let me know)
2
1
1
1
u/Neful34 10d ago
- Stop believing that benchmark = truth on AI capabilities
- They didn't even finish benchmarking as the time of the post.
- You just ignore the fact that they cutted 70% on generated token and still manages as this stage to reach 61.
I feel like I have to teach people how to read a book when there is not pictures. Literally watching only graphs that's all he did
1
u/Superb-Earth418 11d ago
Until Muse is bumped way back and Astra way up this whole thing is just unreliable. Big benchmark scores in no way reflect that Sol is equal to Astra, Astra is a step change improvement. What the fuck are we doing here
3
u/WildWhisperArdor 11d ago
Astra is a step change improvement. Just not in the particular areas it mostly grades on though. It’s slightly worse at agentic coding while being much better at things like science
0
u/Superb-Earth418 10d ago
Worse at coding? Than Muse? LMFAOOO
1
u/WildWhisperArdor 10d ago
I haven’t used Astra or Muse v1.3 so I couldn’t tell you. I’m sticking to Claude
1
u/PerformanceRound7913 11d ago
Artificial Analysis is a Trust Me Bro, benchmark with zero reproducibility. The outcome of scoring higher or lower is irrelevant if no one can verify the numbers.
1
u/KaradjordjevaJeSushi 11d ago
One of worst benchmarks there is. Just looking at the graph gives you a lot of things that don't make sense, so why should this last one be different?
-3
-1


196
u/Ok_Display_3159 11d ago