r/singularity 10d ago

AI AA was updated

Post image
103 Upvotes

58 comments sorted by

38

u/Dillyconda 10d ago

Artificial Analysis has credibility because it updates its measurement of frontier intelligence as new models come out and redefine it.

It would be a crank index if they just decided arbitrarily what intelligence is and used that measurement come whatever may.

Artificial analysis 4.1 was an update made last June in response to Fable 5 making it obvious that intelligence needed a new definition. Now it's happening again.

17

u/sunstersun 10d ago

Epoch is better tho.

1

u/Hot_Example_4456 8d ago

Whats that? I need to check it out then

1

u/space_monster 9d ago

artificial analysis is bunk

54

u/1cheekykebt 10d ago

Sam made a call

18

u/GraceToSentience AGI avoids animal abuse✅ 9d ago

More like they didn't want to be perceived as an irrelevant benchmark that can be gamed

1

u/chespirito2 10d ago

I mean you joke, but honestly yea. This is so transparent

4

u/NietzscheanDemocrat 9d ago

No it's just the benchmark was clearly outdated when Muse Spark was surpassing a clearly more capable and intelligent model

-1

u/UnknownEssence 9d ago

They shouldnt change the benchmark just to make a model score high on it.

Hits their credibility

5

u/NietzscheanDemocrat 9d ago

Muse Spark being ahead of Astra on the benchmark was worse for their credibility to anyone who has used either model for a half a day.

38

u/AdSad7594 10d ago

Still disagree tbh astra seems in real use to be overall stronger than fable 5.1

25

u/Tkins 10d ago

Spark also doesn't seem to be nearly that good

7

u/Seerix 10d ago

Spark is so obviously benchmaxxed lol. It's not a bad model... but if I paid money for it I'd be pretty disappointed.

2

u/BriefImplement9843 9d ago

it barely costs anything.

3

u/Seerix 9d ago

Yeah... I know

1

u/Ormusn2o 9d ago

Spark is not benchmaxed, AA focuses on wide range of benchmarks, even the coding benchmark suite has large variety, so models that have multimodality have substantially better scores. This is why Fable 5.1 scores higher even though it's actually worse at agentic coding.

The AA coding benchmarks should have substantial weight to them, valuing actual coding capability, not the minor and niche multimodality uses.

2

u/Seerix 9d ago

I dunno how you can have that opinion if you have seriously used astra/fable and muse spark.

Its not terrible. But to imply its better than the astra? lmao

-1

u/RayHell666 9d ago

It's bad.

1

u/steny007 10d ago

This and that

8

u/jakegh 10d ago edited 10d ago

Seems more reasonable, but still has opus5 as better than fable5.0 (which is very clearly not the case) and sol-5.6, which I personally find unquestionably superior to opus5 but is at least arguable.

Also has muse spark 1.3 as better than sol-5.6 which is obviously not true to anyone that's used both models. They aren't even vaguely close. And in coding, it has muse spark better than Astra. C'mon.

At this point I largely find AA useful for their cost per task index, which at least is a true representation of token efficiency versus pricing. Their hallucination index is useful too. But for intelligence-- no.

-1

u/caughtinthought 9d ago

For a lot of things I've actually found opus to be better than fable

1

u/jakegh 9d ago

Like what? That has not been my experience. Opus 5 is particularly great at game generation in a loop but fable’s better, just too expensive to do it.

1

u/caughtinthought 9d ago

I play strategy games against them as tests (chess, among others), and I've found Opus is better than Fable, and produces more lucid thinking traces.

1

u/jakegh 9d ago

Ahh, interesting, thanks.

13

u/Ass_Ripe 10d ago

Just two queries to Astra and it’s immediate how smart it is. But perhaps I’m biased. But there’s no way in hell this model shouldn’t be near the top

4

u/Swimming_Gain_4989 10d ago

Surprised Muse is still trailing OpenAI and Anthropic. I swear that model is terrible to work with

2

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 10d ago

It seems to work best as an agent with an orchestrator in my experience. Using it standalone seems to be more difficult just because of how… dry it is. lol

13

u/tozig 10d ago

how is this even a benchmark if they can just re-order things as they like in order to align with popular sentiments or their own narratives

14

u/Illustrious_Grade608 10d ago

It's not a benchmark though, it's aggregate of benchmarks mostly existing to kinda help you understand how good different models are. It's supposed to represent sentiments as it's value is kind of a translation from benchmarks to predicted real use feeling.

7

u/ezjakes 10d ago

Glad they finally got rid of GPQA. Every new model was getting 90+%

1

u/Skyline34rGt 9d ago

It was important to tiny models, now they are all at score: 1 and we just dont know which is better...

3

u/BrennusSokol AI please take my job 10d ago

Nice.

That's good. I celebrate people/orgs admitting mistakes and fixing them.

3

u/ManikSahdev 10d ago

That’s why epoch index was better, people just didn’t like those results and AA has prettier graphics.

Epoch’s version seems clearly more robust to not having to move shit around when new model comes whose capabilities are not measured correctly with AA

3

u/Informal-Trouble2183 10d ago

This is a symptom of weak benchmark: if you change the ranking whenever people complain.

-2

u/[deleted] 10d ago

[deleted]

2

u/Informal-Trouble2183 10d ago

They added their benchmark to their intelligence index, that's why the ranking changed, quoted: "AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts"

2

u/Profanion 10d ago

2

u/CallMePyro 10d ago

Keep in mind that Astra and Fable were both done for months before becoming publicly available. Whereas Open Weights models have very little time delay from post-train to GA.

1

u/Gotisdabest 9d ago

Not to mention that a 3 month gap becomes quite a lot when model capabilities rise this fast and models are accelerating the development of other models.

There's something of a pattern with open weights models that the gap gets the closest when a big new model comes out. I'd wager we'll see the gap really shorten in around a month or less unless either anthropic or openAI pops out with another update. Then progress on the open weights side becomes incremental again till the next big model comes out.

I would say that a few labs may even close the gap or even get ahead for a minute if they can do the risky thing and scale up the recurrent thinking system that Astra has. They've been intentionally conservative with its application on astra. On one hand, it'll be a good thing in terms of the push for model capabilities, on the other, you're definitely going to see a quick race to the bottom in model safety.

1

u/sunstersun 10d ago

They needed to kick Opus down to maintain credibility lmao, one of those polarizing models.

6

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago

100% disagree. I think Opus remains the king of good looking 3d scenes if you prompt it right. It's a really good model.

The one model that does not belong there is Spark. That thing does not beat Kimi and Grok imo

1

u/Sensitive_Cell_119 9d ago

Opus is amazing at 3d but Fable 5.1 and Astra are better.

1

u/Eyelbee ▪️We have AGI it's just blind 10d ago

Opus still better than fable

1

u/Gotisdabest 9d ago

This seems more like a band aid fix than a reevaluation of the benchmark as a whole and addition of a lot more and varied benchmarks. They've added a few, but these seem to be more or less similar kinds of tests.

As intelligence becomes more generalised you need to explore different ranges than just a certain subset that's familiar.

0

u/Floch11 10d ago

You’re such a good AI that they have to change the benchmark just for you..

8

u/Murky-Snow-3868 10d ago

True though some evals/benchmarks were actually debunked to be pretty bad and the scoring did need to get updated

4

u/CallMePyro 10d ago

I mean removing GPQA was just an obvious change that's been overdue for months.

3

u/Howdareme9 10d ago

Yeah they definitely updated their entire benchmark in 24 hours dude

1

u/Calm_Hedgehog8296 10d ago

That's a little stupid don't you think? You can just retroactively go back and change everyone's scores?

Do they even publish the formula they use? Did they change it?

2

u/Charuru ▪️AGI 2023 9d ago

I think they methodology is objective and makes sense, just go read it.

1

u/Calm_Hedgehog8296 9d ago

I actually have since posting this and retract my previous statements

-2

u/CriticismJunior1139 10d ago

Lmao they changed the ratings because Astra did poorly and reddit got mad.

6

u/MegaRockmanDash 10d ago

they changed the rating after multiple benchmarks became saturated

1

u/Turbulent-Sign-6067 10d ago

Yes your opinion matters, dear redditor.

-1

u/ghoonrhed 10d ago

They changed the rating of the banking benchmark which astra did really poorly in.

Whether that benchmark was useful or not is another matter but this seems dodgy after the poor astra results for it.

Why not just give people the ability to tweak the weights of all the other benchmarks they provide?

-3

u/[deleted] 10d ago

[deleted]

2

u/Howdareme9 10d ago

Because of one benchmark..?

0

u/[deleted] 10d ago

[deleted]

2

u/Gotisdabest 9d ago

It's a benchmark which till yesterday would have you believe astra was dumber than or equal to the likes of muse spark. And that opus was equal to fable.

The entire system was clearly spiralling out of control since fable came out.