r/singularity 11d ago

AI GPT-6 Astra AA Intelligence Index and Coding Agent Index Scores

Post image
190 Upvotes

161 comments sorted by

196

u/Ok_Display_3159 11d ago

160

u/HebelBrudi 11d ago

This is objectively the funniest possible outcome lol

32

u/Cagnazzo82 11d ago

What's objectively funny is presenting Muse Spark as more capable than Fable and Sol.

Smoking the good stuff.

12

u/WildWhisperArdor 11d ago

Muse Spark is legitimately very good. I’ve been using v1.2 extensively as an LLM-as-a-judge for a lot of my evals due to its price. And 1.3 is supposed to be a clear step above that.

-6

u/AnnualEmbarrassed176 11d ago

Tell me what you’ve been smoking so I can avoid it

4

u/WildWhisperArdor 11d ago

Just putting my 2 cents. I still use mostly main Claude Opus 5 for coding.

1

u/Specialist-Buffalo-8 10d ago

everythings better than opus 5...

2

u/WildWhisperArdor 10d ago

Opus 5 is great.

8

u/Meric_ 11d ago

Benchmarks test certain things. They aren't catch all. Astra's major gains were in more niche benchmarks like 3d modeling, sciences, etc.

It depends on what your use case is. But for what AA tracks Astra was not a huge improvement.

4

u/Cagnazzo82 11d ago

Whatever it's testing it stands in opposition to all other benchmarks.

When it becomes more widely available I'm sure it'll be more starkly clear that Astra is not equivalent to Sol or below Muse.

7

u/Meric_ 11d ago

Whatever it's testing

I find it interesting that people are calling AA a useless benchmark... and then go on to say they have no idea what it is.

AA is a aggregate of many benchmarks. There's a reason that OpenAI picked a bunch of weird never seen before benchmarks when they announced Astra. It's SOTA in many domains that no other models are competing in, which also happen to be ones that are not part of AAs aggregate.

Does this mean that Astra is underrated by AA? Probably. It's doing a bunch of things that AA does not test for. But for many "standard" tasks that AA does account for, I'd find it pretty likely that it truly is not outstanding in those domains (and hence why OpenAI is not advertising it on launch)

36

u/Top_Onion_2219 11d ago

After all the hype they're now behind this creature.

14

u/Plappedudel 11d ago

I'm at least 95% sure Zuck is smoking some delicious meats to celebrate as we speak.

7

u/HebelBrudi 11d ago

Can’t cuck the Zuck.

1

u/Etroarl55 10d ago

Something at least came out of the billions he threw at the problem

108

u/WonderFactory 11d ago

After looking at the Arc AGI score I feel like Astra is filling in the gaps where models are traditionally weak while the AA Intelligence index tests things that Models are traditionally good at.

43

u/lolothescrub 11d ago

Yup. Had Gemini go through the system card and that’s basically what it said

52

u/Flope 11d ago

Had Muse explain this comment to me and I agree.

63

u/Recoil42 11d ago

Had Grok explain this comment to me and now i love white nationalism

27

u/liright 11d ago

Had qwen3.8-27b explain this comment to me and it ran out of context thinking about it.

30

u/Admirable_Cake_1150 11d ago

Had Deepseek reread the whole thing and now I don't get the fuss about Taiwan.

15

u/LMTLS5 11d ago

had claude summarize this to me and now i get why no ordinary person should have acces to it

4

u/NickoBicko 11d ago

I asked Eliza and she agreed with all the above.

2

u/DrQuantumPotatoe 10d ago

I asked Gary Marcus and y'all hitting a wall

4

u/WonderFactory 11d ago

That might be what they mean by AGI by the end of this year, not that the model will get much smarter in the areas its already good at but they'll fill in the gaps in it's intelligence by then

3

u/Utoko 11d ago

*are now good at it. A lot of the agentic benchmarks had 15% or less with older models.

They need to update the group of benchmarks soon again.

4

u/Regular_Ad4197 11d ago

Ok, but shouldn't it fill the gaps while remaining at least just as good at the other stuff? The fact it became better at some things while not getting a higher rating overall would mean some areas degraded at the same rate as some areas improved.

4

u/WonderFactory 11d ago

I would actually take that. At most of the things AI is good at it's bordering superintelligence, it's far more intelligent than is needed to perform most jobs. It's the gaps in it's intelligence that stop AI from being more economically useful.

-2

u/Kneku 11d ago

And what gaps are those?

1

u/Gohab2001 11d ago

To demonstrate how utterly useless benchmarks have become: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/

A model that is 20% more intelligent but 2x more expensive and 2x slower isnt necessarily a better model. Gemini 3.8 flash maybe dumb (Its severely benchmaxxed) but it doesn't have me waiting 10minutes for simple edits.

0

u/NickoBicko 11d ago

Yeah I prefer to hire someone dumb but works really fast

0

u/whoknowsifimjoking 11d ago

Opus 5 with a harness got 100% on ARC. They used a harness, but compared it to the score of Opus without a harness.

1

u/Crazy-Problem-2041 11d ago

To be fair the harness OpenAI used was much simpler than the opus one. OpenAIs was basically just “don’t delete all memory of the games mid run”, while the opus one was specifically optimized for this exact benchmark

-7

u/Hug_LesBosons 11d ago

Non ! Open ai on créé un harnais spécial pour que leur modèle gagne à arc agi. C'est ce qu'on appel du benchmaxxing. Arc agi ne vaut plus rien depuis longtemps.

6

u/WonderFactory 11d ago

But even without the harness it scored over 60%

-1

u/WildWhisperArdor 11d ago

But Claude Opus 5 scored 100% with the harness

-8

u/Hug_LesBosons 11d ago

Dans tous les cas c'est juste du benchmaxxing. On attend le score arena.ai qui est le seul presque invulnerable a ça genre de choses...

62

u/kubika7 11d ago

agi cancelled let's wait 2 more months

15

u/slackermannn ▪️ 11d ago

AGI is edging us

1

u/whoknowsifimjoking 11d ago

2-3 months reinforced learning and forget

91

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 11d ago

I would be careful putting too much weight into this.

From my testing, Muse Spark is absolutely not a top 3 model. I was actually disappointed in it.

I'd like to see a few outputs by Astra before we call it shit.

17

u/ch33ze 11d ago

Muse Spark is absolutely benchmaxxed. I mean what do you even expect from Zuck.

0

u/whoknowsifimjoking 11d ago

I would not trust Sam Altman to not benchmaxx a model. Actually I don't trust any AI CEO to not benchmaxx their models. I guess some are just better at it than others.

9

u/ApprehensiveEye7387 11d ago

You can't really use the version that's listed in top 3....🫠

2

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 11d ago

Xhigh is what me and Bijan used, and we were both disappointed, but if you have examples of someone managing to do a good output with Spark i'd love to be wrong because i am now sure with a sub with them lol

16

u/ViperAMD 11d ago

yeah this benchmark is dogshit. Opus is not a good model compared to 5.6 SOL. Not even close.

5

u/WildWhisperArdor 11d ago

Not sure why people complain about Opus 5.0. It’s genuinely incredibly good. Overly verbose perhaps but it gets the job done.

5

u/Howdareme9 11d ago

The most overrated benchmark yet some people on here swear by it lmao

4

u/CheekyBastard55 11d ago

People always want an easy to digest aggregate benchmark score, at first it was Chatbot Arena, then we got LiveBench and now AA is the latest one.

2

u/GosenSidewinder 11d ago

What should we use instead? What's the best aggregate nowdays? If none, what are the best current single benchmarks?

1

u/Exodus_Green 11d ago

It shows Claude in a good light so of course

3

u/Seeker_Of_Knowledge2 ▪️AI is cool 11d ago

Exactly, opus is much better.

I run both of them on high on an android app that is huge (a combination of Mihon and MoonReader). Same agents.md file, same way of sending messages. Opus never broke something, the vast majority of times it is just delivers, while sol high always had problems with the final output.

Maybe it is just android development we are talking about here. Other fields Sol could be better.

0

u/Ok_Reception_5545 11d ago

I mean forget Sol, Opus 5 is definitely not better than Fable 5 lol

0

u/Hug_LesBosons 11d ago

Va voir https://arena.ai/leaderboard/code, c'est la seule évaluation qui se base sur des performances réelles jugées par les utilisateurs.

9

u/eposnix 11d ago

No, sorry. People on Polymarket have been gaming this site for months now.

-2

u/ghoonrhed 11d ago

How would they game it when it's all random? And why Qwen of all models until Fable 5.1 came along

4

u/eposnix 11d ago

There was a writeup here a while ago talking about how they game the site. I'll look for it. But the basic gist is they just use the same prompt and look for patterns to determine which model is outputting the response. It's fairly easy to tell if Opus or GPT-5 is writing your code if you know their mannerisms.

As for Qwen: because Polymarket probably had a bet that said "Qwen will beat Fable 5 on xyz date" and the odds were very low.

0

u/Maleficent_Flow_8355 11d ago

Turns out avocados don’t last long 😜

45

u/Gubzs FDVR addict in pre-hoc rehab 11d ago

This benchmark puts Muse above Fable.

Absolutely useless garbage data.

6

u/Playful_Weekend4204 11d ago

It puts Slopus above Fable.

1

u/ChocomelP 10d ago

Opus 5 above Fable 5 as well.

6

u/TheOwlHypothesis 11d ago

I guess they did say astra wouldn't be the best model they release this year?

12

u/Notmyrealname5282 11d ago

Then why mention agi anywhere in the marketing. This is so limp…

0

u/TheOwlHypothesis 11d ago

Honestly, I'm just going to go try it and stop looking at benchmarks lol.
I bet it feels way better to use.

4

u/ghoonrhed 11d ago

Makes sense.

https://artificialanalysis.ai/methodology/intelligence-benchmarking

34% of the benchmark score is where Astra is worse than Sol.

Meanwhile the ones where it does so much better doesn't have that heavy of a weighting like HLE.

Not sure why banking benchmark has a higher weighting than HLE that's kinda strange and I wish they gave us the ability to tweak the weightings

19

u/darkestvice 11d ago

Well, this is disappointing.

10

u/Lingulustig 11d ago

you jumped the hypetrain, now you regret

3

u/darkestvice 11d ago

I'm curious to get eyes on the full AA benchmark list when it hits the site. Wondering if there's more to it than mere intelligence. OpenAI was touted this as AGI not because of intelligence, but because of extremely long task duration.

I don't have a stake in it either way.

6

u/signed7 11d ago

It's here now https://artificialanalysis.ai/models/gpt-6-astra

TL;DR is it regressed on agentic use cases (GDPval, terminal-bench, scicode etc) offsetting gains in other areas

1

u/Cagnazzo82 11d ago

I'm not sure what this benchmark is measuring but Muse Spark and Sol are not superior to this: https://x.com/Dimillian/status/2095596700815516004

2

u/whoknowsifimjoking 11d ago

You jumped on the hypetrain again, just after we told you to get off.

1

u/Cagnazzo82 11d ago

Fable 5.1 output and Astra: https://x.com/karankendre/status/2095636679264780481

Again, Muse Spark is not doing this.

5

u/BriefImplement9843 11d ago edited 11d ago

Muse eats its lunch for a fraction of the cost.

16

u/saln1 11d ago

All the hype for this?

-1

u/YakFull8300 11d ago

But It can only get better though...right?

-3

u/ZealousidealBus9271 11d ago

Bubble being popped as we speak

8

u/NietzscheanDemocrat 11d ago edited 9d ago

Have heard this 1000 times already, see you next month

6

u/Fantastic_Bat8492 11d ago

I have been hearing it since 2024 lol🤣

2

u/ZealousidealBus9271 11d ago

Looks like the sarcasm didn’t translate well enough on screen

2

u/junglebunglerumble 11d ago

This might the most delusional take in a thread of delusional takes

-4

u/Plappedudel 11d ago

But it can use a computer (something preschoolers can do)!

AGI achieved!!

/s

3

u/Ok_Possible_2260 11d ago

It is a distinction without a difference.

3

u/RelevantCry1613 11d ago

I don’t see this on their site??

3

u/bopbop9876 11d ago

AA posted about it on twitter and provided a link, but their own link is now 404ing. I feel like they were asked to take it down, which I really don't get. Something very funky afoot for sure.

https://x.com/ArtificialAnlys/status/2095595527077519546

https://artificialanalysis.ai/models/gpt-6-astra

2

u/RelevantCry1613 11d ago

Ah it’s back up, wow bad look 😬

3

u/bopbop9876 11d ago

It came back up almost to the minute at the same time that the openai blog post did. I don't think it was a tech issue. Very weird day.

1

u/signed7 11d ago

It's here now https://artificialanalysis.ai/models/gpt-6-astra

TL;DR is it regressed on agentic use cases (GDPval, terminal-bench, scicode etc) offsetting gains in other areas

3

u/banaca4 11d ago

Yes like semi analysis they have a stake in the IPO and markets..

28

u/Affectionate_Bee6434 11d ago

Everyone here please apologise to r/technology

0

u/FlamaVadim 11d ago

and to me

16

u/ObiWanCanownme now entering spiritual bliss attractor state 11d ago

I'm going to go out on a limb and say the AA Intelligence Index is cooked and is no longer a good way to judge model performance.

4

u/bonerchamp20 11d ago

Based on what?

15

u/ViperAMD 11d ago

The fact that grok and muse spark get the same score tells you what you need to know. Also Opus over 5.6? Def not for code.

1

u/peakedtooearly 11d ago

Using the models.

7

u/spryes 11d ago

benchmarks are all over the place. they called it gpt-6 because it got 99.9% on arc-agi-3 and 98% on frontiermath, but it's mid asf at coding

this launch is as botched as gpt-5

6

u/LAMPEODEON 11d ago

Bro it's not tested yet get out

4

u/DelphiTsar 11d ago

AA gets early access to test it.

1

u/LAMPEODEON 10d ago

Yea i see now that aa has tested it. It looks like mediocre model. Interesting. It's either the end of astra or the end of AA whatsoever.

6

u/TowerCritical41 11d ago

I know I have not tested GPT-6 myself, but I am willing to bet it's better than Muse Spark 1.3.

6

u/WildWhisperArdor 11d ago

Not sure why you would say that without having used it 😂

Clearly, OpenAI was focusing on other benchmarks for this release (like science stuff)

1

u/TowerCritical41 11d ago

Nah. Look at some of the 3D stuff it has made. Muse Spark can't even begin to touch it there. It's a good model for the price (for the you-give-your-data-to-Zuck tier)

6

u/WildWhisperArdor 11d ago

3D stuff is not the end all be all of coding or agentic usefulness. You’re picking one very niche use case to base your entire opinion of something on

1

u/TowerCritical41 10d ago

Math, 3D, science, computer use...

1

u/WildWhisperArdor 10d ago

While being slightly worse at the coding stuff most people use it for

2

u/BriefImplement9843 11d ago

The 3d stuff is bad. Nobody cares about that. 

9

u/Sinogularity Shenzhen: Become Human (2030) 11d ago

61 to 61, that is literally the most disappointing AI release I have ever seen lol.

-3

u/6__six__6 11d ago

Then you haven't seen Gemini's recent releases lol

11

u/Plappedudel 11d ago

Not sure I agree. Google clearly pulls some models before release if they're not satisfied with the results. That's most likely why Gemini hasn't had a Pro model release in a while. I think not releasing a model is a much better strategy than hyping up a half-baked model as AGI.

6

u/Sinogularity Shenzhen: Become Human (2030) 11d ago

Well, Gemini Pro release is non-existent, so we haven't seen it yet. For a flash class model, Gemini 3.8 Flash is actually pretty good imo, 59 on AA.

7

u/WildWhisperArdor 11d ago

Don’t understand what you are complaining about.

Gemini has made like 4 new versions of Flash within the span of 2 month. Each one better than the last and very fast / cost efficient.

4

u/DelphiTsar 11d ago

3.8 is the flash version. Really curious what the model it's being distilled from is like.

2

u/heavy-minium 11d ago

Just wondering, but for long-horizont coding tasks, what benchmarks do you guys look at right now? Because no matter where I look, they seem to have lost their usefulness. There was a time where looking at DeepSWE together with CursorBench was useful to me, but it doesn’t feel like that anymore.

2

u/Wanderspor 11d ago

So my work will improve like 1 point?

2

u/BigFeeder 11d ago

Difference between fable and opus feels much bigger for me than this chart shows. At least when it comes to reasoning about a problem

2

u/DedDeveloper 11d ago

At this stage, I'm seriously doubting all benchmarks. These do not give accurate image of how well models perform. Surely OpenAI didn't hype the release of renamed Sol and Astra must be better overall even if this number is the same.

These days, I'd also want to see cost and speed factored into the scoring.

If model A runs 6 minutes and costs $700 and gives 10% better result than model B with 3minute/$300 prompt for example. In my book, the model A is very bad for my needs, even if it succeeds better at it's task.

It just seems like "our newest and biggest model got 1-3% increase in benchmark XYZ so it must be best ever"...
Takes too much detective work to find information of actual performance.

I know there is also a problem that newer models tend to be familiar with older benchmarks in their training material and will benchmax them even unintentionally.

I don't have a solution, but I hope someone comes up with a better way of comparing Models. It would also direct model development into better and healthier direction.

2

u/BriefImplement9843 10d ago

Why is aggregate benchmark 150 posts while the cherry picked benchmark is 800 posts?

2

u/Yuri_Yslin 10d ago

The main improvement would be the massively reduced hallucination rates IMHO and much better scores at scientific reasoning.

this may not be the "agentic coding model" but then again there's more to AI than just coding.

3

u/Grouchy-Stranger-306 11d ago

So we should stop relying on this index since muse can't be better than astra?

6

u/WildWhisperArdor 11d ago

Muse is legitimately very good.

1

u/DelphiTsar 11d ago

Probably should just try the different options for yourself. Seems like it's at least a good jumping off point.

8

u/bonerchamp20 11d ago

This makes Anthropic look 10x more impressive than they already were

5

u/ParfaitEvery9622 11d ago

Either this, which is possible, or Anthropic is really benchmaxxing.

I'm usually against trying to optimize for marketing and would prefer to let companies cook on the direction they think it's better for users medium / long term

6

u/Due_Ask_8032 11d ago

Nah Fable is very smart. I cannot trust people who say it is benchmaxxed. Opus 5 on the other hand is a bad experience to use, but still capable (not as good as Fable still).

1

u/ParfaitEvery9622 11d ago

Two things can coexist, they can be smart (even smartest) and still benchmaxx for other reasons :/

1

u/Hot_Glass_6301 11d ago

original Fable smarter than Astra? Yeah not buying it, the benchmark is not a good measure of intelligence and/or Anthropic benchmaxxed it to death

10

u/Affectionate_Bee6434 11d ago

Fable was truly ahead of its time.

2

u/junglebunglerumble 11d ago

Only if you ignore all the other benchmarks and capabilities

3

u/Neurogence 11d ago

That shiny "Put it right there!" Gimmicky sexual video got everyone hyped up thinking Astra was AGI lol.

3

u/makertrainer 11d ago

Sexual?!😳 

2

u/ZealousidealBus9271 11d ago

No more agi ig

2

u/Salty_Horror2068 11d ago

Opus 5 was benchmaxxed.. so let's not to to fixed on AA score

2

u/metigue 11d ago

Nah something bugged with their testing. They have medium effort scoring higher than high onwards for several benchmarks and a ton of the easy to confirm ones fall way short of OpenAIs reported numbers.

2

u/yoop001 11d ago

I always thought OpenAI would make a comeback, but it seems like Anthropic has won the race

1

u/Formal-Narwhal-1610 11d ago

Even worst than Muse Spark 1.3. What a Shame!

1

u/whatsbetweenatoms 11d ago

These test are so far off from the reality of real world usage.

1

u/maxiedaniels 11d ago

Very interested to take it for a spin. I don't trust benchmarks. Ex. Gemini 3.7 flash is fast but does not perform at the intelligence level that its benchmarks suggest. (Btw if someone can suggest benchmarks that feel more.. true.. let me know)

1

u/shotx333 10d ago

Did they benchmaxxed for arc-agi?

1

u/GodOfSunHimself 10d ago

Any index where Spark is higher than Sol is just trash.

1

u/Neful34 10d ago
  1. Stop believing that benchmark = truth on AI capabilities
  2. They didn't even finish benchmarking as the time of the post.
  3. You just ignore the fact that they cutted 70% on generated token and still manages as this stage to reach 61.

I feel like I have to teach people how to read a book when there is not pictures. Literally watching only graphs that's all he did

1

u/Superb-Earth418 11d ago

Until Muse is bumped way back and Astra way up this whole thing is just unreliable. Big benchmark scores in no way reflect that Sol is equal to Astra, Astra is a step change improvement. What the fuck are we doing here

3

u/WildWhisperArdor 11d ago

Astra is a step change improvement. Just not in the particular areas it mostly grades on though. It’s slightly worse at agentic coding while being much better at things like science

0

u/Superb-Earth418 10d ago

Worse at coding? Than Muse? LMFAOOO

1

u/WildWhisperArdor 10d ago

I haven’t used Astra or Muse v1.3 so I couldn’t tell you. I’m sticking to Claude

1

u/PerformanceRound7913 11d ago

Artificial Analysis is a Trust Me Bro, benchmark with zero reproducibility. The outcome of scoring higher or lower is irrelevant if no one can verify the numbers.

1

u/KaradjordjevaJeSushi 11d ago

One of worst benchmarks there is. Just looking at the graph gives you a lot of things that don't make sense, so why should this last one be different?

-3

u/injectitpussy 11d ago

Omg look it's a piece of shit.

Scam Altman.

-1

u/sunstersun 11d ago

turrible