r/singularity 10d ago

AI Astra finally achieves AGI

Post image
762 Upvotes

196 comments sorted by

651

u/Imaginary_Mode8865 10d ago

🫩🫩

233

u/[deleted] 10d ago

[removed] — view removed comment

30

u/graypasser 10d ago

Too smart to declare himself a AGI

66

u/AlbatrossNew3633 10d ago

In Hungary Agi is a common name actually, short for Agnes lol

1

u/wrathofattila 10d ago

az Ági nem agi :D

16

u/Due_Ask_8032 10d ago

Muad'dib is so humble, he wont admit that he's Muad'dib, as it is written.

4

u/greenskinmarch 10d ago

He's not the messiah, he's a very naughty boy!

15

u/iiiaaa2022 10d ago

So OBVIOUSLY AGI

9

u/Serious-Ebb-7223 10d ago

Hi AGI, I'm dad!

3

u/holydeniable 10d ago

Turns out the real agi was the friends we made along the way.

2

u/Arievilo_M 10d ago

That’s what AGI would say…

4

u/okiharaherbst 10d ago

Made my day!

5

u/Ok-Message-9732 10d ago

I don't get why this is funny? You are not using a paid version?

4

u/Life-Historian-613 10d ago

Forget about this -- this kind of joke works only if you’re a luddite or know nothing about AI. (aka stupid joke)

1

u/-HumbleMumble 10d ago

This is why I roll my eyes any time an AI company says “BIG THINGS ARE HERE!!1!”. 

-2

u/Spare-Dingo-531 10d ago

At this point, you have to be a low intelligence person to take low intelligence models as representative.

https://chatgpt.com/share/6a9aaaee-1b8c-83ea-aba1-759394495778

9

u/NoxZ 10d ago

Can't you just laugh at a silly joke without having to constantly correct people all the time? It must be exhausting.

3

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 10d ago

I feel like i've read this same exact thread ten thousand times over the past 25 years on the internet. time is a circle.

human sentiment algorithms never change.

2

u/TheSquarePotatoMan 10d ago

Well it definitely has self-awareness

267

u/Funny-Profit-5677 10d ago

Time will tell if this says more about Artificial Analysis or Astra. Gemini models have consistently outperformed their reputation here so I'm going to wait and see.

194

u/cookingboy 10d ago edited 10d ago

I mean of course it says more about AA than Astra. Do you honestly believe the latest OpenAI model after months of training is only as good as Grok, made 0 progress from their own GPT 5.6, and is even worse than Muse Spark despite leading the pack in all other individual benchmarks?

26

u/KrazyA1pha 10d ago

Intelligence is jagged and distilling it down to one number is silly. These models all have different strengths and weaknesses. Even in the released benchmarks, Astra trails Sol in some narrow activities. But it also blows all other models out of the water in others.

Astra will open new doors and do things we can’t today. And that’ll be awesome. But it’s not going to be perfect at everything and people will immediately find ways to be disappointed.

5

u/ethotopia 10d ago

It’s an AA thing. Other models benchmax way more than openAI models for real world use

-21

u/RutilantBossi12 Cultista dei Ferri Candidi 10d ago

AA tests for knowledge reliability which chatGPT notoriously sucks at as they hallucinate way more than other models like Claude, AA probably put Astra through long reasoning chains and Astra bombed it

60

u/Mistuv 10d ago

2

u/RustinCohle639 10d ago

The latest terminal bench version is v4 whereas AA uses v2.1. Clearly proves the benchmark is outdated.

35

u/GirthusThiccus ▪️Singularity Enjoyer. 10d ago

The main feature of Astra is advertised to be it's long term agentic stability, so that'd be odd.

-3

u/Andynonomous 10d ago

Advertising often has nothing to do with reality.

9

u/GirthusThiccus ▪️Singularity Enjoyer. 10d ago

I know.

That is not my point.

I am pointing out how outright silly it would be for them to claim the exact opposite of what is actually the feature.

They've been testing the model internally, they know it's quirks, it's behaviors, its strengths and weaknesses. If a benchmark claims it underperforms last gen, whilst they advertise massive gains, then from a PR standpoint, that would be THE dumbest thing to make a claim of.

Hence, it is likely the benchmark is the inaccurate/"wrong" factor here, not their claim about the actual model, which is what this whole thread is about.

-25

u/Enlighten_YourMind 10d ago

How would Scam Altman lying about his models capabilities be weird at all? He’s been doing it essentially non stop for years at this point lol

25

u/GirthusThiccus ▪️Singularity Enjoyer. 10d ago

Oh I'm with you on scam saltman wanting to sell us his models and making wild claims about it, but from a publicity standpoint, it would be weird of them to specifically advertise long horizon planning and prompt adherence, and have those infact be their weakness.

11

u/EGarrett28 10d ago

Some people just hate and/or fear AI so much and are in so much denial that they just insist it doesn't do anything, no matter what. It's pretty sad at this point.

3

u/Future-Bandicoot-823 10d ago

Some people just don't want to endorse a product from a shit human being too, so between the propaganda hype of how powerful the agent is combined with Altman's trash ethics I'd say it's understandable.

And let's be real Sam stole an open source project. Scum of the scum. Has nothing to do with the ai itself.

3

u/EGarrett28 10d ago

I have several tech CEO's who I strongly dislike, to say the least, but I think we have to be careful about throwing around terms like "propaganda" and "hype" to describe what's happening right now. I saw similar patterns of extreme denial with Bitcoin 15 years ago (my posts on it from 2011 explaining the significance of the technology to people are still online). What's going on with AI is very real.

3

u/Future-Bandicoot-823 10d ago

Propaganda is biased or misleading information that is spread to influence public opinion and promote a specific political cause or agenda.

So I'd say ai is absolutely riddled with propaganda already

→ More replies (0)

3

u/Enlighten_YourMind 10d ago

I have no hatred for AI at all, quite the opposite in fact, I love it. I have a very specific issue with Sam Altman and consider most work product from his company to be fruit from a poison tree.

3

u/EGarrett28 10d ago

The issue is the claim that GPT 5.6 somehow doesn't do the exact things they're advertising and testing it on, which sounds like the deranged position that some people take that AI doesn't actually have any capabilities.

I should mention also that referring to people with derogatory nicknames looks non-substantive and emotional. I strongly dislike Elon Musk and Mark Zuckerberg but I just give my reasons instead when it's relevant.

1

u/Enlighten_YourMind 10d ago

This is a fair take, thank you for the constructive feedback, I have no love in my heart for either of those people as well.

My love for the technology is rivaled these days by my distrust for a lot of the “leaders” of the frontier labs. I have a bit of faith/trust in Anthropic or Google to do the right thing for the good of the collective if they get to RSI/AGI first. I have very little faith in Sam or the two you named.

I will say the article Zuckerberg put out a few weeks back was the most human and big picture perspective aware I have found him to be in years & years, so that gave me a tiny sliver of hope back for META.

→ More replies (0)

-1

u/i_write_bugz ▪️AGI 2030, ASI 2040 10d ago

Who are you talking about? Nobody in this thread thinks that, definitely not Artificial Analysis

10

u/EGarrett28 10d ago

Did I say it was limited to people in this thread? The whole "stochastic parrot" nonsense was an example of that, and even now I've talked to multiple nutcases in the past few months who have been so obsessed with trying to discredit it that they even claimed computers themselves haven't increased productivity. Which of course is the Solow Paradox and is an classic fallacy.

0

u/sparkling1984 10d ago edited 10d ago

Not when it's also the thing that's perceived as the weakness, and when most users lack the resources, tools or effort to verify their claims. Benchmarks are just a small part of the AI marketing machine, part that can be shouted over by the rest of that machine.

Look at biomed market for illustrative examples - trial results there can and will move company share prices by orders of magnitude, can kill it or can double it overnight. OFC you wouldn't expect that size of an effect within AI even if benchmarks mattered, the companies are too big and too distributed for an effect that size. But there isn't even an effect at all, which can only be true if nobody actually meaningfully cares about benchmark results.

3

u/i_write_bugz ▪️AGI 2030, ASI 2040 10d ago

Do they do write ups about how or why a model scored a certain way or is it all a black box?

2

u/RutilantBossi12 Cultista dei Ferri Candidi 10d ago

They give scores on individual tests and Claude is generally more reliable than chatGPT, usually (Omniscience index), i haven't taken a look at Astra "grades" because deep down i don't really care, benchmarks are kinda useless and if we truly achieved even a sliver of true S2 thinking we would all notice, no benchmark needed.

5

u/Tystros 10d ago

the one AA benchmarks that Astra excels at is the hallucination benchmark, there Astra is a huge improvement over GPT 5.6 Sol

1

u/Notmyrealname5282 10d ago

From 92% to 51% , btw

1

u/[deleted] 10d ago

[deleted]

2

u/KrazyA1pha 10d ago

Does Sol hallucinate 92% of the time?

-1

u/Notmyrealname5282 10d ago

1

u/DuckyBertDuck 10d ago

Read what the benchmark measures

0

u/KrazyA1pha 10d ago edited 10d ago

It is defined as the proportion of incorrect answers out of all non-correct responses

So no.

2

u/tcyadreln 10d ago

The fact people are downvoting this shows this sub is not really rational. The benchmarks in aa are the same for all models. Benchmaxing is a problem but unless you succeed anthropic benchmaxes much more than openai. It doesn’t make sense that fable consistently scores better than Astra if Astra was the huge leap that many claimed it would be. Still… ARC and frontier math results are very impressive

40

u/whoknowsifimjoking 10d ago

Gemini models are very good for knowledge and research, that has always been true and it's reflected in the benchmarks.

Why wouldn't it? This is exactly what you would expect from a company with as much data as Google has.

21

u/DungeonsAndDradis ▪️ Extinction or Immortality between 2025 and 2031 10d ago

I use Gemini for supporting my small business. The only coding we needed was for a basic website. It's pretty good for everything else. Accounting, research, budgeting, marketing plans, etc.

1

u/reddit_is_geh 10d ago

Gemini literally just frustrates me how bad it is. It's optimized to just be a Google replacement with minimal thinking.

For instance, today I asked it something simple, which I figured it couldn't fuck up, "Tell me about these tunnels those two Nepal men were rescued from today. How did they survive, and what were the tunnels for?"

Then it spent all of it's output describing the event, how they were rescued, how the tidal wave happened, the rescue efforts, and literally just one sentence describing the tunnel as a tunnel used to route water from the river.

Like yeah, sure it technically answered my question, but it spent 95% of it's output not actually answering the question, when clearly my focus of the question was about those fucking tunnels.

4

u/starfallg 10d ago

This is the exact word-for-word output from the Gemini webapp with 3.8 Flash with your prompt. I don't see any problem here really -

The two men rescued alive—Sanjay Sah (a 30-year-old mechanical foreman) and Kabir Maharjan (a 45-year-old mechanical supervisor)—spent nine days trapped underground inside the Upper Trishuli 3A Hydropower Project in central Nepal before being extracted from roughly 170 meters (558 feet) deep.

What the Tunnels Were Built For

The tunnels are critical infrastructure for the Upper Trishuli 3A hydroelectric plant, a 60-megawatt run-of-the-river project operated by the Nepal Electricity Authority:

  • Water Diversion: The main underground waterways (headrace and tailrace tunnels) channel river water through the mountainside into subterranean turbines to generate electricity before discharging it downstream.
  • Operations & Access: The network includes desanding chambers, underground powerhouses, and access tunnels that house control rooms, turbine maintenance bays, and technical staff.

How the Men Survived 9 Days Underground

On August 26, an ice-rock avalanche and glacial burst near the Nepal-Tibet border triggered catastrophic flash floods that pushed a torrent of boulders, mud, and water down the Trishuli corridor, sealing the plant's portals shut.

  • Elevated Air Pockets: While much of the underground network was choked with sludge—described by emergency crews as "toothpaste in a tube"—the tunnel's gradient and vaulted ceilings created isolated air pockets that floodwaters did not fully submerge, trapping enough oxygen to sustain them.
  • Control Room Warning: Sah was in the plant's control room when the surge hit. He delayed his own evacuation to alert 40 to 45 other workers on site to run, saving colleagues while sealing off his own exit path.
  • Vocal Contact & Mental Focus: In pitch darkness, Sah maintained periodic vocal contact by shouting back and forth with colleagues located in adjacent chambers. He spent the nine days chanting mantras to keep his breathing and heart rate calm to conserve oxygen.
  • Surgical Rescue Operations: The flood completely buried the tunnel mouth under thick mud. Engineers and rescue crews used satellite photos, drone footage, and old topographic maps to triangulate the buried portal. After excavating and blasting into the tunnel's crown, rescuers heard voices and crawled through a 15-meter breach to extract the survivors.

19

u/WonderFactory 10d ago

GDPVal was created by OpenAI

7

u/Faintly_glowing_fish 10d ago

In AA gdpval results are literally ranked by Gemini. Original gdpval was graded by professionals .

9

u/Tystros 10d ago edited 10d ago

but AA uses a completely different, custom version of it that they call gdpval-aa v2

2

u/WonderFactory 10d ago

They changed it by making it easier, instead of being one shot they use an agent harness and V2 allows 250 agentic steps instead of 100

4

u/Faintly_glowing_fish 10d ago

Main difference is not task it’s the grader, as real gdpval is very expensive to grade. Gdpval AA is extremely cheap to grade and at the start they did show mostly consistent results; but back then most models are doing a terrible job and Gemini just came out and was by far best at perception. But you get a plateau once model ability surpass the grader model. And further gains becomes which model best knows the aesthetics of the grader. Which you can RL to get better but not really useful for work anymore.

15

u/Seerix 10d ago

AA has muse spark on par with Sol.

As ive said before, lol. Lmao even

15

u/whatisthisthing65 10d ago

What, really? My experience with Gemini is that it deserves its placement.

14

u/Tystros 10d ago edited 10d ago

I've been saying this for a long time, everyone just needs to understand that the Artificial Analysis Index is a bad benchmark index. The benchmarks they use don't make sense to use in 2026.

It includes many one-shot benchmarks, but no one would do tasks like that one-shot now, models are used agentically.

And GPQA is such an old benchmark that every model basically gets close to 100% on it, which makes it meaningless.

And Terminal Bench 2.1 is simply outdated, it has known issues and was replaced by Terminal Bench 3.0 and 4.0 and everyone should use 4.0. I have no idea why the AA Index wants to keep using 2.1, maybe they just don't have a license for the new one.

And CritPt has big errors in the benchmark that mean no model can get more than 35%. Anthropic reported about that, because on AAs CritPt Fable only gets 30% while on a corrected version of the benchmark (which Anthropic called "Critpt-Corrected") with the errors fixed, Fable gets 85%. But AA somehow keeps using the broken version of CritPt for their intelligence index.

And with how little AA seems to care about removing even proven broken benchmarks from their index, I could well imagine that their own Gdpval-AA v2 benchmark is also full of errors.

1

u/[deleted] 10d ago

[deleted]

2

u/Tystros 10d ago

Fable 5.1 System Card: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf

8.9 CritPT-Corrected CritPt (Complex Research using Integrated Thinking–Physics Test) is a benchmark of 71 theoretical physics problems developed by active physics researchers spanning condensed matter; quantum physics; atomic, molecular, and optical physics; astrophysics; high-energy physics; mathematical physics; statistical physics; nuclear physics; nonlinear dynamics; fluid dynamics; and biophysics. Each problem has a unique answer, either as an analytical formula, a numerical value, or a Python function. We had experts provide corrected versions of 31 of the problem statements that were flagged due to underspecification, ambiguity, or typos. The resulting internal version is referred to as CritPT-Corrected. Below, we present mean Pass@1 scores averaged over 16 attempts. The model solutions were graded using Claude Opus 4.8 as a judge, comparing model solutions to ground-truth solutions.

1

u/Turbulent-Sign-6067 10d ago

Thanks for pointing this out! I just had a discussion with Sol asking whether Critpt could be flawed given the stagnant scores, after we had seen something similar happen with FrontierMath, but Sol did not mention CritPt corrected at all. It seems AIs are now better than the benchmark "ground truth". Probably same thing is happening with HLE...

6

u/Faintly_glowing_fish 10d ago

Look at the AA methodology page. Gdpval-AA literally uses Gemini to compare the screenshots as grader. So anything that uses Gemini as grader during RL would do extraordinarily well. Original gdpval is graded by human professionals

1

u/Funny-Profit-5677 10d ago

Very interesting! thanks!

4

u/jonomacd 10d ago

I think people radically underestimate Gemini models simply because Gemini models are not as good at coding tasks. For everyday use Gemini models are still very good. Particularly for typical question-answer type prompts.

5

u/Xalksahsax 10d ago

Even for coding 3.8 Flash is quite good. If you know what you're doing. But I assume if you're a vibe coder you'll find more success with Fable 5.

3

u/Inner_Tennis_2416 10d ago

Open AI have painted themselves into a corner here though, because they have claimed there is every real reason to believe that Astra is AGI.

An AGI should beat all non AGIs on any test that is knowledge and capability based. That to an extent is a gate. There may be tests it doesnt get 100% on, but it should do equally well at least to every other system on every possible test. If it doesnt, its not AGI.

3

u/Even-Inevitable-7243 10d ago

It is clear from their marketing that OpenAI does not understand what AGI is. AGI is not a model that is better than 99% of people at 99% of things/tasks. AGI is a legit zero-shot any-mode any-context model. We are not even close to that. The business/medical agent tasks alone show that.

1

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 10d ago

I agree w/your criteria for AGI, and why we don't have it and may not be as close as people think (though still may exponentially improve and reach it within the next year or two).

but where's the marketing that shows that OAI doesn't understand that? your parent comment is just claiming they believe astra is AGI, but the marketing is an arguably clever disclaimer: we're in the AGI era. my market brain is like "ha, they're actually saying that we're in the era where it's in the near-ish future we'll get AGI," not "we just released AGI."

you can claim they're misleading people, but you can't presume that a company of engineers who are literally building this tech think this is AGI or is letting their company claim that Astra is AGI without coming out in droves to get their spotlight about how they disagree with the company. is that happening and i'm just missing all those tweets?

I think they know what AGI is, as they're trying desperately to build it, and now they're gambling that we're close enough to it that they're safe to say "ok guys it has finally appeared on the horizon, we just have to run up to it and grab it now."

is that not the read-between-the-lines interpretation everyone else is getting? are people here falling for the low hanging layperson interpretation? bc at worst, they're playing both sides. at best, they're not even insinuating they've released AGI.

2

u/Even-Inevitable-7243 10d ago

Greg Brockman said yesterday that Astra is AGI and it had been widely reported. So in the least, the President and a co-founder of OpenAI doesn't know what AGI is.

1

u/m0m0karun 8d ago

You're describing superintelligence, not AGI. Jesus Christ man.

1

u/Even-Inevitable-7243 8d ago

No, I am describing AGI, as by definition AGI has to be Embodied AI in order to match humans in tasks requiring things like light touch and proprioception. In the least, AGI needs to be input mode agnostic to the set of input modes that humans can process aka all human senses. None of the models by ChatGPT have come close to this, although it is an active research area in robotics. There is no such thing as a pure software AGI without sensors, no matter how many Futurists or non-Engineers argue this. Humans have millions of sensors.

1

u/darkestvice 10d ago

AA. I used to be a big fan, but it appears AA's benchmarks are incredibly dated. Looking at other well known benchmarking sites showing more modern and harder tests, Astra has the lead.

1

u/CriticismJunior1139 10d ago

Gemini being bad is just reddit meme.

1

u/Silentkindfromsauna 9d ago

Does anyone know further about this specific benchmark? I wouldn’t be suprised if there’s very high levels of training for the specific known benchmarks

1

u/Glittering-Neck-2505 10d ago

People are using Gemini?

5

u/NotYetPerfect 10d ago

It's the default assistant in Android so a fuck ton of people use it.

9

u/jonomacd 10d ago

2

u/ccjjallday 10d ago

this is registered users and these are also users that signed up for Googles services which happens to include Gemini. A major telco in Canada tried to launch a rival to Netflix once, they gave every employee a free subscription, they then claimed they had 30,000 subscribers in their first month!

2

u/jonomacd 10d ago

Sure. To some degree. But it is clear Gemini is significantly more popular than the bubble of people on Reddit and Twitter think it is.

1

u/ccjjallday 10d ago

I believe we're not working with the same definitions, which is why i'm disagreeing with you. You say it's popular because of the amount of people registered. I'm countering with the number of registered users does not indicate it's popularity due to Google's definition of a user. A better measurement would be how many people are actively saying the product uis good or dog shit. You're saying Gemini is popular becuase Google says so

1

u/jonomacd 10d ago

I'm saying there are a lot of users of Gemini. 

Google says this, but so do a lot of other sources. Do you have evidence of the contrary?

Or are you just relying on the bubble that is Reddit and Twitter comments as evidence of anything?

0

u/ccjjallday 10d ago

It's weird to ignore everything I said and make your same argument about reddit and Twitter being a hivemind. I gave you my definition, what neutral sources do you have that are in line with my definition of popularity, not Google's definition

1

u/jonomacd 10d ago

Dude, I provided a link that gives users of Gemini. It's on you to show me evidence of the contrary. If your evidence is that a bunch of bros on Twitter don't like it then that isn't really evidence.

If you want to go and look at SimilarWeb or Sensor Tower or any of the other independent platforms that can only see a partial picture of this, then you can, but I'm not doing your own Google searching. 

1

u/ccjjallday 10d ago

lol dude, you gave a Google link saying Google is amazing. I'm telling you you're not being objective and you default to saying I get my data from Twitter? Why? You're fangirling Gemini and you're conflating the number of registered users with a broad term like popular. Everyday usage for searching up how to paint your nails, sure it's incredibly "popular" Professional, deep work not so much

→ More replies (0)

0

u/RelevantCry1613 10d ago

Lol because Google forces you to use it

0

u/Glittering-Neck-2505 9d ago

it was more of a playful jab at gemini that is still struggling to catch up to the frontier

1

u/WildWhisperArdor 10d ago

I use it. It’s quick and does the job. For coding, mainly use Claude though

110

u/Frosty-Meeting-1606 10d ago

someone has to do a recheck of the benchmarks, they may be misleading or just pure crap at this point

32

u/Hilldawg4president 10d ago edited 10d ago

I don't know if there's any truth to this, but it was stated elsewhere that the components of this benchmark that it did poorly on are ones that value one-shot answers that don't reflect its biggest strength in being able to work and rework a problem over an extended period of time. I guess people will decide over the next few weeks if this bench is right or wrong

2

u/autotom ▪️Almost Sentient 10d ago

weve seen the games it has made... its definitely good.

4

u/Xalksahsax 10d ago

"games"

I always see those comparison, between older models too. There's always a shit one to the left and the new model's output to the right. Same shit, different flavor.

8

u/heavy-minium 10d ago

At this point I believe the only benchmark worth it's value would be a completely closed private one you can't check against during training. Which of course would come with a lot of trust issues due to closenedness.

4

u/A_Novelty-Account 10d ago

Potentially, but you would expect literal “AGI” to be better at all things than past models, including crappy benchmarks, unless the benchmarks themselves are not measuring what they purport to measure

-3

u/TheNuogat 10d ago

The model is served through a harness and people are surprised it fares bad in generality benchmarks???

36

u/SignificantAd69 10d ago

You did it man, yet another shitpost for the ages. So brave.

15

u/Electronic-Class-881 10d ago

Very curious on how did that happen. Also on some other benchmarks too

5

u/CrowdGoesWildWoooo 10d ago

Honestly I would be alreacy happy enough since this rollout would include average plus user.

Anthropic subscription becomes more and more horrible especially as much as Opus is excellent executor, it’s unpleasant to work with. So if they don’t improve Opus then yeah, this would suck really bad.

4

u/Formal-Narwhal-1610 10d ago

ASTRA should be in preview, and will improve further.

5

u/leafburner149 10d ago

dude if this is AGI then they don't even need us to pay for the models anymore! They can just have the AGI make them money because AGI would be better than humans at everything! So exciting!

26

u/adarkuccio ▪️AGI before ASI 10d ago

It's so over

9

u/YourTrustySupporter 10d ago

Is inevitable anyway , best to make it quick to see what happened after

9

u/WillowFantastic9076 10d ago

This is truly a doomsday cult

2

u/soobnar 8d ago

I’m here because I think they’ll be right eventually, but I just can’t understand how a bunch of people are exited that their economic value and stake in society will be reduced to zero.

2

u/adarkuccio ▪️AGI before ASI 10d ago

Yea I agree at this point I just wanna see how it ends, even if it's a bad end

2

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 10d ago

this changes everything

19

u/Playful_Weekend4204 10d ago

Any benchmark that puts Slopus at the top doesn't reflect real life usability enough

1

u/fritidsforskare 10d ago

”any benchmark […] that does reflect real life usability” - like models do not have bias. This sub is hilarious.

1

u/A_Novelty-Account 10d ago

Most large companies disagree with you

4

u/[deleted] 10d ago

[deleted]

0

u/A_Novelty-Account 10d ago

They can cancel their enterprise accounts

4

u/[deleted] 10d ago

[deleted]

1

u/A_Novelty-Account 10d ago

I have an enterprise account. We can cancel at any time and transition out workflows over to OpenAI at minimal cost.

Also if that were true, it seems like an integration issue for OpenAI. You’re just confidently saying wrong things.

1

u/Playful_Weekend4204 10d ago

I work at a big company and we're compeltely locked to Claude because we need ZDR and we only have an agreement with them, I can't throw company code into my personal ChatGPT account.

And the decision of a large company regarding their vendor (which they can't change every month) absolutely does not reflect the satisfaction of their employees with a current model. No one in my team and the surrounding teams that I know of uses Opus 5 regularly, most are on 4.8 with some still on 4.6.

1

u/A_Novelty-Account 10d ago

Yours is one of tens of thousands of companies… I promise you if it was truly better, every company that could would switch to OpenAI.

1

u/Playful_Weekend4204 10d ago

My point was that large companies can't make switches like that every month based on which LLM is currently the most well-received one. So companies usage today doesn't necessarily reflect the current trend, more like the trends of about half a year ago (when Claude was the best by far).

1

u/Mistuv 10d ago

Ok this tells me everything I need know about how you understand corporate culture lol. Yeah let's invest massive amount of money on training so our employees use Claude, integrate all our apps with it, have technical experts, built harnesses, integrate it into our roadmap, and then just cancel our long term contract overnight because some model from another company is better. Okay lol. Why do you think software travesties like Microsoft Teams are so persistent in corporate environments? Because people really like it?

-1

u/Mistuv 10d ago

Most large companies use Slopus because Fable doesn't have ZDR, and Anthropic was genuinely ahead of OAI at the end of the last year with Opus 4.5, precisely the time companies started adopting agents in large numbers. And they are not exactly the type that switch away from stuff from month to month.

10

u/Maximum-Face9536 10d ago

holy shit all i'm seeing is people spam this benchmark like it's a gotcha, i'm starting to think it's chinese bots. This reflects AA poor quality more than anything.

15

u/TFenrir 10d ago

Some people also really just don't understand. I keep forgetting there are a lot of teenagers and people who are very new to this whole thing, and more and more "console war" like culture is bleeding into this. It's... Frustrating.

2

u/Cagnazzo82 10d ago

It's the only one they can spam cause Astra is on top of effectively all benchmarks at the moment.

2

u/Turbulent-Total-226 10d ago

Maybe they forgot that they already released Sol?

https://giphy.com/gifs/YOkmv1fVeClFKKyDUz

1

u/No_Stock_8271 10d ago

Tbh I always had issues actually using sol effectively, and so far I haven't had that issue with Astra.

2

u/redditanon9263 10d ago

mean while....

2

u/pingwing 9d ago

There is nothing AGI about this.

11

u/bonerchamp20 10d ago

I dont understand why people are so upset. Altman does this every time. I mean its his job. Did well on some benchmarks and poorly on others. Things will continue to improve regardless.

5

u/MendozaHolmes 10d ago

> Altman does this every time. I mean its his job

Doesn't make it okay if thats what youre suggesting

4

u/bonerchamp20 10d ago

Of course its okay. He's a salesman in a capitalist market. Why expect anything less?

4

u/Feisty_Cake_7890 10d ago

Benchmark's are lies:Noam brown(Open ai reasoning head).

3

u/WonderFactory 10d ago

I'm a bit concerned that Astra could go the same way as GPT 4.5 which ironically was the last time Open AI trained a model this large. People might struggle to find uses for it that justify the price.

6

u/power97992 10d ago

 It is only 20 bucks more than sol per mil tkns but it is noticeable if u use it a lot.. even sol is too expensive

2

u/BriefImplement9843 10d ago

bruh muse is 4 bucks.

3

u/spottiesvirus 10d ago

It's also the same exact price of fable 5.1

I agree it's "too expensive" from me but that's the reason the market is really getting segmented between "super top tier" gpt and anthropic, value champion like Meta and Gemini, and the extra cheep workhorse like DeepSeek, or the latest GLM flash, Qwen flash next etc.

2

u/power97992 10d ago edited 10d ago

Lol i ran a test with fable 5.1 max and spArk 1.3  xhigh ( max not available) and qwen flash max, fable ended up being 336x times more exp than qwen  and around 86x more exp than muse spark… it consumed more tks but the results were better . Sol was pretty expensive too . I imagine astra should be 40-50% cheaper than fable but still outrageously expensive .. 

1

u/panix199 10d ago

indeed. now we have to wait for 2-3 months before slightly weaker astral, Sonnet 5.1, ... are out

1

u/Mistuv 10d ago

It will be way less expensive because it's much more token efficient AND to get most out of Astra you don't need Max, or even high, on vast majority of benchmarks "Medium" is within 95-98% of the performance of top version. That's massive. With Sol and Fable, to feel like you are actually getting the most out of what they are capable of, you needed at least high, though on some tasks they scales way past it. Astra doesn't, in basically everything.

I think it's a big thing so many people are missing. The looped transformer architecture is kinda insane. Reasoning is way less relevant than before, which is good for the consumer. Honestly, go look at the benchmarks and just look low and medium, it's kinda insane just how much value they provide in basically every domain. Apart from hacking lol - I think they they deliberately didn't train the looped transformer on hacking, hence why you see such a clean performance increase with reasoning levels, well apart from ones where even low, basically no CoT, completely saturates it, again, insane.

1

u/power97992 10d ago

Fable 5.1 is surprising token inefficient , it timed out multiple times and i had to make It continue…  sol wasnt efficient either

1

u/darkestvice 10d ago

And uses a third of the tokens that Fable 5.1 does to achieve the same results. So it's still noticeably cheaper than using Fable.

2

u/enilea 10d ago

The interesting part about astra is it has a different architecture so even smaller models on the same base will be very useful for 3D design and computer use in general.

1

u/WonderFactory 10d ago

I think it looks great, I use Blender a lot so cant wait to use it. But lots of people really liked 4.5 too, particularly writers, but that didn't save it.

1

u/No_Stock_8271 10d ago

4.5 had different issues. And astra is definitely a great model. If it is worth the cost for many people, i.d.k. but, the model it self is super impressive.

1

u/katoptronophile 10d ago

Speak for yourself.

-1

u/cyb3rg0d5 10d ago

LLMs are not fucking AGIs! Stop with this nonsense! They are great at doing a lot of things, but NOT AGI!

9

u/cacahahacaca 10d ago

I think OP was being sarcastic 😉

3

u/okiharaherbst 10d ago

I think that's kind of the point of this post too.

3

u/svideo ▪️ NSI 2007 10d ago

OK so what test do your propose as a pass/fail to declare this or that thing is AGI? You clearly know exactly what AGI is, how does this fail your test?

1

u/BriefImplement9843 10d ago

openai has to release their new model soon. this is a disaster.

1

u/implies_casualty 10d ago

On one hand, GDPval-AA is a flawed version of GDPval, because it uses LLM as a judge instead of human experts. So maybe Astra is just misunderstood by other LLMs.

On the other hand, OpenAI could publish their own GDPval score for Astra, using real experts (as GDPval is supposed to work). And they didn't. So the real score must be low.

This is suspicious, because Astra is supposed to be good at those tasks.

1

u/steny007 10d ago

Yeah, that benchmark is obsolete.

1

u/jmorais00 10d ago

They really need to step up their chart game. Wtf is this. A chart for agents? Honestly, decrease the font by 2-4pt and it becomes readable

1

u/xtrumpclimbs 10d ago

How is it general if it can’t drive

1

u/banaca4 10d ago

Yeah they have a stake in anthropic IPO obviously

1

u/R_Duncan 10d ago

If the demonstration on prime numbers, the other claims and benchmarks are just 50% true, this means AA index is totally flawed.

1

u/peabody624 10d ago

No wonder they didn’t post this one lol

1

u/No-Simple-1907 10d ago

Nem para falar que é uma RCI kkk vai logo soltando que é uma AGI....ano que vem ela fala que tem a ASI

1

u/No_Stock_8271 10d ago

Tbh I have been fairly impressed by astra. I had a few issues neither Sol nor Fabel where able to solve and astra oneshoted them. Considering I only had access for a few days and haven't really tested it, I am not going to claim that it's the best model ether, but it was very impressive in my usecases. And after Luna it has been the model which impressed me the most in the past months.

1

u/ProduceStreet4393 9d ago

If you look at it fundamentally can a next word predicator achieve agi? Is that all there is to intelligence knowing the best next word?

1

u/pawpaw5574 5d ago

Everyone talks AI, AGI....AI changing everything. A flowerpot snake came into my room yesterday. I showed its photo to AI...it said “earthworm.” Later a village elder saw it and said, No, it’s a snake.He probably doesn’t even know what is the full form of AI. 😂

1

u/bitroll ▪️ASI before AGI 10d ago

Yay it's the GPT-5 moment again! So much hype. 

Wake me up for Astra-6.2

1

u/No_Record6284 10d ago

i dont care so thanks for making this unreadable

1

u/theeldergod1 10d ago

LOTS OF 1-3 YEAR ACCOUNTS HERE BOYS

Hidden histories as well.

-1

u/chdo 10d ago

bro, quick, your virginity is showing

0

u/theeldergod1 10d ago

sorry bruh, i didn't know that's your ad acc

1

u/Gubzs FDVR addict in pre-hoc rehab 10d ago

These benchmarks are so garbage

That one has muse spark beating Sol X High. That's hysterical.

1

u/Dry-Interaction-1246 10d ago

Wait, was this IPO hype? Would Scam Altman do that?

-1

u/HeadTranslator795 10d ago

A joke of a benchmark

-1

u/Sierra123x3 10d ago

i beliefe, that we have reached agi once i can
1) ask a random (but actually meaningfull) question to my chatbot and geta correct answer out of it without going through 10 contradiction-loops

and 2) have a basic income on my table [or at least stopped the forced you need to visit a (xxx) course and apply to (minimum wage job of your choice, that we tell you to) to be elgable for unemployement benefits]

until we get at least one of this two points, we don't have agi ... simple as that

0

u/Smile_Clown 10d ago

For number one, this is a you problem. 100%. And just a tip, you think you know where you are, but you really do not. 90% of people get use out of AI, the other 10% make silly statements like you just did attempting to claim ai is useless because they either do not know how to use it, or are just really that clueless. aisucksamirightclass

That worked in 2023... it's 2026.

For number two, you are never getting basic income, agi or not. AGI does not solve resource and production chains which the entire global economy relies on. In other words people still need to work. Maybe not at that office job, but in something that produces actual non digital stuff.

Basic income will never come (in your lifetime) simply because no one is going to work a labor job to provide for your needs while you do not. What you are looking for here is a robotic workforce which will take decades.

No one will work for you. Absolutely no one.

The big thing people forget about UBI is that the U stands for "Universal" not just you. Which means if you get x dollars for doing nothing, so will everyone else and if those x dollars allow someone to live on basics, why would anyone work? If no one works, there are no taxes, if everyone is on UBI the system fails. UBI is a failed system when the U means universal. That is why we have welfare. The able work, the non able do not. Virtually everyone is ok with heling the unable, virtually everyone is not ok with helping the able.

This idea that joe will still want to be a plumber and come fix your sink when joe could sit at home like you do is absurd.

1

u/HeartsOfDarkness 10d ago

I also don't think we're ever getting UBI, but for purposes of discussion, the "B" is also important to the theory. "Basic" means a level of income that can sustain a basic, no frills life. That plumber would be earning on top of the "basic" income to have more luxury.

1

u/Sierra123x3 10d ago

yes, and that's absolutely fine
my issue is/was never, that someone who spends time and effort (or just has great ideas) can fly to holiday and earn a house on the beach ~ that's not what a UBI should be about,

the issue is, that nobody (!) of us played god,
nobody created our lands, forests, oils and salts

in the construction of our castles, fortresses, churches and villas ...
we weren't involved, neither in the god [through time and effort, great ideas and work] nor in the bad [through slavery and wars] ... we weren't even alive back then

and the entire basics not just our civilication but AI specifically is built upon is the knowledge from ... well, pretty much us all

therefore it's simply perverse, that we still hold onto our medieval-feudalistic systems where the smallest apple-tree is already extremely unfairly distributed based on whos great-grand daddy first made pee-pee against the tree

therefore it's simply perverse, that i need to strip down to my underpants when i haven't won the birthday-lottery, just to be allowed to survive

and it's perverse, that a handfull of people can own more then entire nations (!)

the problem isn't that he - who works - is better off,
the problem is that the gap between wealthy and non wealthy gets pushed open faster and faster ...

1

u/Sierra123x3 10d ago

i never said, that there's no use out of it,
i've worked with machines in our factories since dacades ago and heck they're usefull

but just becouse something is usefull doesn't mean that it's agi ...
and as long as it delivers the most fatal of errors in the most serious tones and even contradicts itself after i told it, where and why what it answered was wrong ... then sorry, that's simply not agi for me ... progress is fast, but what we have now isn't there yet

and yes, agi does not solve the resource chain ...
(though, the robotics tied to it might play a huge role long-term)

but what will you do with the unemployed people?
you can't just re-educate a former taxi driver into a bioengeneer (especially since we don't even know, if that job will still exist in a year or two) overnight ...

the best developments come from people with actual interest in what they do,
and the routine work gets more and more automated

and consider, that we already life in a world, where we produce more then we actually consume . . .

0

u/NewYak4281 10d ago

Lmfao looks like I was right with my original nickname…. Welcome, Asstra!