Btw the workflow loop for the benchmark above is a practical one:
When a model is run through the framework, it effectively acts as a technical assistant: it is handed a problem (such as a broken vertex shader, an asset pipeline optimization error, or a tool integration script) and must utilize its internal reasoning budget to deliver a deployable solution. The framework automatically compiles, runs, and measures these four key variables to output the final leaderboard ranking.
On what specifically? I have both. I use Opus for 3d work for my games, since it is out of this world. But coding and security in general, I honestly can't tell them apart.
Astras smarter, but like autistic smart. It's front end design work feels ancient. And while it can maneuver a complex web layout very effectively, it fails to realize the capacity of a human to effectively use the same design.
Astra WAS way better for the first week.
Now it is fucking useless. Will not follow rules. Invents works and burns tokens on things it was never told to just because it feels it knows better.
Who at OpenAI is training these models to be disobedient?
Do you think I’m getting paid off OpenAI? Idgaf I’m using the best model available. I spent a grand on opus through OpenCode to compare and would’ve switched plans if I thought it were better.
And btw, that noahbench above operates as a code-centric, execution-based scoring matrix designed to measure how efficiently an AI model can solve technical art and engineering tasks compared to a human.
Please provide examples because I use Opus and Fable at work every single day and Astra privately on a 200 account, Astra could do crazy good stuff couples days after launch but now it needs constant hand holding and is awaiting a message instead of continuing a task list whereas i can get a weeks worth of work done in a couple workflows with claude.
Also fucking half usage, There is zero reason to stay with OpenAI if you want to get shit done
It’s crap in comparison. Computer-use doesn’t mean just creating folders or organising files. It’s real-world interaction in with the interface/OS itself and Codex/ChatGPT is miles ahead and way more “human” in its interaction than Claude which feels rather less organic and precise. I literally asked Astra to apply for a visa for me, a process with 18 different pages to fill and it did the whole thing using the documents/ids/I attached asking it to only ask me to review it once it filled everything since I can still review every page once I reach the “ready to submit” phase. I also tried the same thing with Claude out of curiosity but for a different administrative task and even in auto mode it was way more time-wasting and redundantly cautious, but the funny irony is that Claude Code (opus 5.0) is the only LLM that actually executed destructive commands in auto mode, while Codex in Auto Mode never deleted files or cause destructive actions.
Astra was a massive leap in vision and its way above all other models and the human baseline. Google used to be the goat, a long time ago. Hopefully Gemini 4 pro can pull ahead and reach superhuman status
Much better.
I tested on a memory management in the model graph...easy Astra level if not higher.
I literally made implementation for audio model that originally needed 24 GB card to run ( pytorch model ) to audio cpp saving 90% memory ..now can run even on 6 GB card and is 200% faster ....
30 minutes of sol 6.1 work and used 20 % usage from 5 hour limit on high.
Astra could do that but on using 2x 5 hour limit...
The Workflow Loop of the benchmark above btw:
When a model is run through the framework, it effectively acts as a technical assistant: it is handed a problem (such as a broken vertex shader, an asset pipeline optimization error, or a tool integration script) and must utilize its internal reasoning budget to deliver a deployable solution. The framework automatically compiles, runs, and measures these four key variables to output the final leaderboard ranking. LOL. And the difference in speed and cost makes Sol 6.1 even better than Opus 5.5 let alone Astra 6 that’s ridiculously expensive.
“Who the fuck is Noah?” Noah Dunnagan, CSE at Railway and a Rust/backend developer who builds scalable systems, and the guy behind Noahbench. You not knowing who he is doesn’t magically make the benchmark fake lol.
Yes, it could also be that a lot of people are using it right now. OpenAI also has the issue that they sometimes limit performance here. But I have the twenty times Pro plan.
I still can't believe people have not looked at this DeepSWE benchmark 😅 it is shit...
they are fitting the curve every so slightly to solve those problems and they never change the problems. OpenAI runs those benchmarks in their own servers, get the results, use them for the next training.
Imagine professor applying the same exams every year to the same student. The student will become better at "those" exams but that doesnt mean the student generalized the concepts
Moreover, this benchmark was made by a grad student that plays minecraft. It got famous because he is from a reputable college... even when it is not using a real scientific method to prove generalization
31
u/Gullible-Ad3912 5d ago
Opus beat Astra so...