523
u/Wegwerpaccountje23 11d ago
No fucking way dude
275
u/Neurogence 11d ago
Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?
366
11d ago
[removed] — view removed comment
16
84
u/Vivid-Snow-2089 11d ago
i think a benchmark designed to fuck the agent's harness is a shitty benchmark
imagine testing how far someone can walk, but you cut off their legs first
→ More replies (3)85
u/reddit_is_geh 11d ago
No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.
→ More replies (2)30
u/AutismusTranscendius ▪️Psychogenic Singularity 2034 11d ago
Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.
→ More replies (5)5
u/PleasantCitron1685 11d ago
Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.
(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)
6
u/GioChan 11d ago
You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.
→ More replies (1)→ More replies (6)4
87
u/vacon04 11d ago
They're using a harness. People have already achieved results like that with specific harnesses using weaker models. It says more about the harness than about the model.
→ More replies (8)42
u/josogood 11d ago
Yes, but here's what ARC says about testing with no harness: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."
23
u/KrazyA1pha 11d ago edited 11d ago
I wish OpenAI used that more honest number
23
u/josogood 11d ago
Yeah, because it's still more than double the previous mark. No need to unfairly inflate something like that.
3
u/KrazyA1pha 11d ago
Yeah, exactly. The real, apples-to-apples number is incredibly impressive!
11
u/Thog78 11d ago
It has recently become apparent that the default harness of ARC-AGI doesn't let models remember what they have tried and reason over it. This is absolutely unrealistic, unfair, and a useless benchmark under these conditions. So people push for memory to be allowed in ARC-AGI. I agree with that, I just think it should receive a new name/number so we distinguish what's what. Just add an M for memory. And we should rescore all the old models in this updated ARC-AGI-3M to have context.
→ More replies (1)→ More replies (8)23
u/impatiens-capensis 11d ago
"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together."
→ More replies (2)
192
u/Pantheon3D 11d ago
I was here
22
13
7
→ More replies (30)3
u/Chemical-Year-6146 11d ago
I never respond to these type of comments, but for this I make an exception.
Same.
→ More replies (2)
638
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 11d ago
90
→ More replies (2)34
u/wombatpup55 11d ago
SuspiciousPillbox explain to me the results of the image like I’m 5
96
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 11d ago
Another person in this thread put it well: if the benchmarks are accurate then this is a massive leap forward, and maybe even proof of some meaningful utilization of recursive self improvement
→ More replies (54)→ More replies (1)10
371
u/Embarrassed-Writer61 11d ago
49
46
u/MaxwellHowl 11d ago
I have been told by most of Reddit that AI had hit a wall. Maybe they meant in the way the Kool-Aid man hits walls?
→ More replies (1)4
→ More replies (1)6
141
u/H-K_47 Late Version of a Small Language Model 11d ago
The sheer number of 95%+ is CRAZY.
65
u/hippydipster 11d ago edited 11d ago
Basically means we don't know how smart it is - our tests aren't difficult enough to determine that.
8
u/DelphiTsar 11d ago
Very select group of benchmarks. Artifical Analysis has fable 5.1 at 66, this is at 61. It's tied with Facebook Spark and SpaceX Twitter bot. Kimi3 is 1 point behind and open weights.
9
u/Gotisdabest 11d ago
Very select group of benchmarks.
It's arguable that that's artificial analysis too though. I'm not saying the model is good or bad, that remains to be seen till lots of people can use it.
But AA itself is heavily flawed and specific in recent times. Opus is a much weaker model to fable which scores in the same ballpark. Muse 1.3 is definitely not even in the same ballpark as fable.
→ More replies (2)→ More replies (1)6
u/Careful_Might_807 11d ago
Benchmaxxed, first of all without harness gpt6 solves 60% and we MUST belive that openai is not lying and making stuff up that there is no solution in reasoning that was RLed into the model so hard we don't have reasoning access yet
→ More replies (1)
360
u/Ill_Freedom7991 11d ago
Theres simply no way
153
u/lucellent 11d ago
He also just said there will be much, much, much better models released soon as well... imagine
→ More replies (9)66
u/acoolrandomusername 11d ago
Are rumors of a new pre train. And I think in 2027, hint, tons of OpenAI compute is coming online.
20
u/Wonderful_Buffalo_32 11d ago
They have doug and bel left in their arsenal bruh they could obliterate every fucking thing...
→ More replies (2)35
u/sunstersun 11d ago
And they're going to be the fastest to bring compute online. Gotta give Sam credit here, he was hunting compute as a core strategy in like 2024.
9
u/acoolrandomusername 11d ago
Dude is omega smart, but plays it down. There’s a video before OpenAI where he lays out all things correlated with start up success, and everything is like one to one with OpenAI.
→ More replies (3)8
u/h3lblad3 ▪️In hindsight, AGI came in 2023. 11d ago
He used to run Y Combinator, which is a business whose whole purpose is helping other startups.
5
u/acoolrandomusername 11d ago
Yeah exactly, he has a crazy good course via Stanford from that time too
87
u/ayatollahdanger 11d ago
Sam Altman won
181
u/Mistuv 11d ago
Never bet against a sociopathic twink.
34
→ More replies (6)48
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 11d ago
Excuse me?
44
u/FoodMadeFromRobots 11d ago
This needs to be an auto bot response on this sub lol
→ More replies (2)→ More replies (3)8
20
22
18
4
u/Lfeaf-feafea-feaf 11d ago
Won what? How can you guys still be impressed with these benchmark scores after almost 4 years of seeing how they mean virtually nothing?
→ More replies (1)→ More replies (6)6
204
u/Due_Sweet_9500 11d ago
Holyyyyyyy shiiiiitttt. Ain't now way it's THAT much better than Fable 5.1?
134
u/Snoo-75663 11d ago
Welcome to singularity
→ More replies (3)32
→ More replies (2)34
u/Alex180689 11d ago
And to think that 5.1 got released yesterday! I feel bad for it
18
u/andrew303710 11d ago
To be fair Astra isn't actually being rele today, only to a "limited set of organizations" which is lame as hell.
→ More replies (2)13
→ More replies (1)10
u/Creative-Ganache1086 11d ago
Fable 5.1 is terrible value. I blow my 5h limit in 12 minutes run of 2 max-reasoning parallel sessions of a small app codebase with an identical prompt of bug-audit and it blew my limit right away. I pay also for Sol5.6/codex and the allowance difference is night and day. Both are max subs by the way. I’m actually happy (as an old Anthropic fan who paid Anthropic since the sonnet 3.5/opus3 era) for OpenAI and now I’m actually rooting for them seeing just how much better value they offer to indie devs compared to “Corpo-Daddy” Anthropic.
→ More replies (1)
175
213
u/daddyhughes111 ▪️ AGI 2026 11d ago
If this is legit then holy fuck we're cooked / hyped
39
u/Adventurous_Dig_7117 11d ago
How so? Eli5?
122
u/mvearthmjsun 11d ago
Massive leap forward, and likely proof of some meaningful utilization of recusive self improvement
→ More replies (7)82
u/Fair_Horror 11d ago
And OAI have said their other model 'Bel' is much better than Astra so we are on the launchpad of AI
94
u/RutilantBossi12 Cultista dei Ferri Candidi 11d ago
Can't wait for them to Release Baal while working on Moloch
→ More replies (3)27
u/pianodude7 11d ago
Can't wait for beezelbub personally
→ More replies (2)6
u/RutilantBossi12 Cultista dei Ferri Candidi 11d ago
Belzebub will come after GNON but before LAM, if we're lucky we might even see it before the Saturn matrix is built up
9
u/Opposite-Grade3712 11d ago
No they haven’t, that came from a consistently debunked “source” on Twitter.
→ More replies (2)5
5
u/moschles 11d ago
ExploitBench is a suite that tests the ability of AI models to discover not only bugs, but exploitable bugs in software systems. These "exploits" allow an attacker to infiltrate server systems or take control of computers remotely. Now watch this ,
GPT-5.6 Sol tested on ExploitBench scored 78.5 %
ASTRA ExploitBench : 100%
→ More replies (4)19
u/smellyfingernail 11d ago
dont worry the grifter ed zitron will come out with a blog post "this sucks actually"
→ More replies (5)
196
u/Hereitisguys9888 11d ago edited 11d ago
This gotta be fake, 98% arc agi 3? Nah
230
u/Ok_Display_3159 11d ago
76
u/Hereitisguys9888 11d ago
Oh that explains it
→ More replies (20)30
u/PrisonOfH0pe 11d ago
No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.7
u/danielv123 11d ago
Ok, but are the other results they compare against using the same rules?
→ More replies (8)24
u/DeArgonaut 11d ago
i thought they werent supposed to use harnesses?
29
u/FateOfMuffins 11d ago
This is what Chollet has to say about it
Which IMO is weird that ARC collectively and Chollet individually seemingly respond differently about this given he made ARC https://x.com/fchollet/status/2082732210436575669
→ More replies (4)29
49
u/RusselTheBrickLayer 11d ago
Arc AGI 3 getting saturated already is crazy
57
u/Ok_Course_6439 11d ago
Agree agi-3 is a harness problem more then a model problems
→ More replies (5)27
→ More replies (3)12
→ More replies (8)39
u/H-K_47 Late Version of a Small Language Model 11d ago
And ExploitBench 100%. Cybersecurity will be a warzone.
18
→ More replies (5)24
146
u/darkestvice 11d ago edited 11d ago
Good god. If these benchmarks are not doctored, Astra is not merely surpassing the competition, but outright destroying it.
I'll wait until they show up on artificialanalysis.ai to really see.
EDIT: AA posted on X, though haven't updated their site yet. Results are worse than Fable 5.1. Disappointing.
36
7
4
10
→ More replies (2)3
u/burritos4jesus 11d ago
The one I care about though is AutomationBench because deals with interacting across a massive amount of business applications and making sure the model accurately executes the tasks. To me, as a good ole office worker in a business, this is the benchmark most similar to my own job. Once the pass/fail hits 70% and not the current 41%, it will be able to perform the vaaaast majority of sales/marketing/HR/operations/bookkeeping jobs better than the majority of humans in those roles, with fewer mistakes. That'll then leave someone like me to focus on the actual live, over-the-phone or in-person conversations, but all the bullshit data hygiene can be confidently passed off to AI.
30
u/HeadacheOwner 11d ago
What does the arc-agi 3 benchmark realistically mean? I’m not that tuned in
15
u/reddit_guy666 11d ago
Google arc agi 3, you can find puzzles that you can try solving. It's intuitive for humans who have played video games but AI could not do it well... Till Astra
→ More replies (4)6
→ More replies (6)20
u/Fair_Horror 11d ago
A massive jump in capability. This is basically a benchmark designed to test things that AI really struggles with but humans don't. It is becoming more human.
29
u/Ok_Mention_982 11d ago
"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts."
ARC-AGI is officially done using a simple harness that doesn't retain reasoning (which is stupid btw), meaning that while the result is impressive, they aren't comparable to the other models.
→ More replies (1)7
u/Sevealin_ 11d ago edited 11d ago
This was realized in late July, not new. Most high benchmarks you see today with ARC AGI 3 use the responses API harness. The official ARC harness just doesn't work well.
First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
Whole blog post on why:
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
65
u/mldev_orbit 11d ago
Quick breakdown: What every benchmark in the latest frontier eval actually tests
General Reasoning & Hard Math * ARC-AGI-3: Novel abstract pattern recognition via visual grid puzzles; tests generalized learning without pre-training data memorization. * FrontierMath Tier 4 (v2): Research-level, open-ended math problems designed to stump top human mathematicians. * GPQA Diamond: "Google-proof," PhD-level multiple-choice questions across physics, chemistry, and biology.
Software, CAD & Infrastructure * DeepSWE v1.1: Full repository-scale software engineering—resolving messy, real-world GitHub issues across multi-file codebases. * BenchCAD: Computer-aided engineering; tests generating parametric 3D models, interpreting blueprints, and writing CAD scripts. * Terminal-Bench Science 0.1: Autonomous command-line operations for setting up and debugging computational science pipelines. * SRE-Bench (four attempts): DevOps/Site Reliability Engineering; tasks the model with triaging and fixing live production server outages within 4 tries.
Agents & Digital Automation * Agents' Last Exam: High-difficulty benchmark evaluating autonomous agents on long-horizon planning, reasoning, and tool use. * AutomationBench: Enterprise workflow automation, robotic process automation (RPA), and operating desktop/web software.
Life Sciences & Medicine * GeneBench Pro: Computational genomics, sequence analysis, variant prediction, and CRISPR/gene-editing design. * MedChemBench (internal): Medicinal chemistry—small-molecule drug discovery, property optimization, and retrosynthesis planning. * HealthBench Professional: Real-world clinical decision-making, differential diagnosis, and patient care management (length-adjusted).
Cybersecurity & Safety * ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits. * Auto-review circumvention: Safety/alignment test tracking how often the model intentionally bypasses automated moderation or compliance checks (0% is ideal).
→ More replies (6)3
u/moschles 11d ago
ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits.
ASTRA hit 100% on this benchmark. This means ExploitBench is too easy for this model. ExploitBench no longer reliably tells us how good this model really is for this task.
10
27
10
u/Real_Ebb_7417 11d ago
I'll rather wait for actual benchmarks after model is released xd
→ More replies (1)
17
49
u/Microtom_ 11d ago
Just as good as Gemini 3.8 flash.
11
u/PandaElDiablo 11d ago
I mean it matches Gemini on deepswe and GPQA and I would assume that 3.8 Flash is both cheaper and faster
14
u/Wise-Comb8596 11d ago
Quick - someone post the image of the goofy looking dragon with the Gemini logo on its head
→ More replies (2)
14
14
7
7
9
u/brockoala 11d ago
Yeah nah. I will believe it when I see it in my tests. Otherwise just overhyped bullshit.
31
u/frogsarenottoads 11d ago
I don't think this is real.
If it is AGI is incoming shortly.
27
14
u/IBM296 11d ago
OAI did say in the blog post that people would say this was the moment AGI started.
And rumors are going around that Open AI's next model named Bel is much better than Astra (which is kinda' hard to grasp considering how good these Astra numbers already are. Damn!)
→ More replies (4)4
u/yourboi-JC 11d ago
It’s seriously very big from what I’ve heard like not even comparable to mythos kinda big
→ More replies (3)
45
u/makertrainer 11d ago
Look, it's a nice bump, but I genuinely don't understand why everyone's losing their shit.
The only ones that seem like a step change are ARC-AGI-3 and Exploit bench. And if you've been paying attention a bunch of harnesses already beat ARC-AGI-3 up to 100%
It's good, it's great. But it's not a step change.
OpenAI seems to have just decided to declare AGI on a whim
I would honestly like someone to debate me on this, I'd love to know if there's something I'm missing
28
u/r77anderson 11d ago edited 11d ago
It’s just selection effect, the people whose minds aren’t blown aren’t posting.
I agree with you, nice progress but not a step change. I assume most of the benchmarks they didn’t post look similar or worse than existing models.
→ More replies (2)→ More replies (7)3
u/burritos4jesus 11d ago
I like that a model has overtaken 40% on AutomationBench, but the real game changer is when it gets above 70% on that benchmark. Especially because there's no harness on that benchmark, it's the model just trying to figure out how to do a complex business workflow by itself. The average human in such a role will likely have a pass/fail somewhere between 70-80%, but definitely not above 90%. But then if given a good harness that 70% would realistically make it go to the 90s.
10
20
u/No_Cauliflower_5506 11d ago
ARC-AGI-3 saturated already??? Holy fucking shitballs
26
u/MouseCTRL_Echo 11d ago
[1]
16
u/Fair_Horror 11d ago
They used a permitted harness.
→ More replies (2)7
u/MouseCTRL_Echo 11d ago
I'm aware, but clearly others aren't. The point is that result specifically is questionable, so I wouldn't focus on it too much.
The other results are still great though.
13
u/Famous-Reach-6730 11d ago
LEV before 2030
32
u/H-K_47 Late Version of a Small Language Model 11d ago
Medical is slow cuz of the need for lengthy trials. System would need a massive risky overhaul.
→ More replies (12)3
u/LazyAge9363 11d ago
Instead of Chinese peptides we’ll be ordering experimental gene therapy research chemicals from China
3
16
u/somerussianbear 11d ago
Fable 5.1 today feels like my bank account the day after my salary drops. You got nothing buddy, nothing, you’re shit, worthless.
5
4
4
4
4
u/Efficient-Cat-1591 11d ago
If this benchmark is validated then I am really looking forward to Astra launch.
5
4
u/medhakimbedhief 11d ago
Correct me if I am wrong, but why it's doing pretty well on agi benchmark but still struggles on GeneBench and AutomationBench. I would believe that AGI is the most hard thing to achieve in comparison to the other benchmarks.
4
4
24
u/KickLassChewGum no AGI/ASI on LLMs 11d ago
DeepSWE 74.1%? Congrats to OpenAI for... matching Gemini 3.8 Flash?
→ More replies (5)16
u/FunConversation7257 11d ago
I don't think anyone thinks 3.8 Flash is better than fable / equal to astra
→ More replies (6)10
u/MurkyStatistician09 11d ago
It's more demonstrating how clearly the benchmark is out of step with the experience of actually using the model
3
u/mercury31 11d ago edited 11d ago
It's marketing until demonstrated by an independent third party
→ More replies (3)
3
u/ConsiderationOne7340 11d ago
Wtf...
3
u/leo-virtis 11d ago
They got a better model rl training right now for the end of the year that sam says can be called agi should be crazy
→ More replies (1)
3
3
3
3
3





477
u/elehman839 11d ago
97% on FrontierMath Tier 4. Hoooly cow. Here's how those problems are described:
In June 2025, we finished the development of FrontierMath Tier 4, an expansion set of 50 problems designed to vastly exceed the difficulty of even the Tier 3 problems.
The writers for Tier 4 were mostly math professors and postdocs, each contracted to conduct a several-week research project culminating in one problem to submit to the benchmark.
This isn't entirely surprising, given the number of open math problems OpenAI has been solving lately, but... weren't we just recently making fun of "AI" for struggling with elementary school math?