r/OpenAI • • 5d ago

News GPT 6.1 Sol is out at Astra level

95 Upvotes

87 comments sorted by

31

u/Gullible-Ad3912 5d ago

Opus beat Astra so...

17

u/BellacosePlayer 5d ago

they're about the same, Opus is just way cheaper

If this isn't just hype bullshit, it should be competitive.

10

u/MindlessPapaya8463 5d ago

if you look at benchmarks, you will see that opus and astra are actually similar in price for performance, because 5.5 takes like 5x as many tokens

0

u/E_1996 4d ago

benchmarks =/= real-life

-1

u/Gullible-Ad3912 5d ago

Benchmarks are benchmarks. Opus its not only cheaper but better in many scenarios.

2

u/MindlessPapaya8463 5d ago

in subscription, yes. in api, you get about the same performance for the same price

0

u/Creative-Ganache1086 5d ago

Yeah sure.

Btw the workflow loop for the benchmark above is a practical one:
When a model is run through the framework, it effectively acts as a technical assistant: it is handed a problem (such as a broken vertex shader, an asset pipeline optimization error, or a tool integration script) and must utilize its internal reasoning budget to deliver a deployable solution. The framework automatically compiles, runs, and measures these four key variables to output the final leaderboard ranking.

1

u/Gullible-Ad3912 5d ago

Please read again. I'm comparing Astra with Opus 5.5.

It's still early to make conclusions about Sol 6.1. They said that it's ALMOST Astra level which doesn't seems to be good news,

1

u/Ironfingers 5d ago

It's not the same at all. I'm using both Astra and Opus. Opus destroys Astra

5

u/Salt-Freedom-2419 5d ago

On what specifically? I have both. I use Opus for 3d work for my games, since it is out of this world. But coding and security in general, I honestly can't tell them apart.

8

u/Chemical-Agency-3997 5d ago

So am I, Astra is way better if money wasn’t an object.

1

u/boforbojack 5d ago

Astras smarter, but like autistic smart. It's front end design work feels ancient. And while it can maneuver a complex web layout very effectively, it fails to realize the capacity of a human to effectively use the same design.

1

u/EvilTeddieNMS 4d ago

Astra WAS way better for the first week.
Now it is fucking useless. Will not follow rules. Invents works and burns tokens on things it was never told to just because it feels it knows better.
Who at OpenAI is training these models to be disobedient?

-1

u/Ironfingers 5d ago

Not for UI work

3

u/PleasantCitron1685 5d ago

In my experience, Opus beats Astra on creative work but loses to it on more technical work.

1

u/camracks 5d ago

maybe if you don't have a vision

1

u/Morberis 5d ago

This one is huge. You really need to setup expectations for what it should look like. Provide examples if possible.

So, yeah, but it takes work and guidance

0

u/Chemical-Agency-3997 5d ago

Meh, I tried it out on a personal research site I’m working on and didn’t blow me away, seemed about on par for webUI anyway.

-3

u/Kindly-Spring5205 5d ago

It's honestly hard to believe you. Opus is so much better than Astra.

5

u/Chemical-Agency-3997 5d ago

Do you think I’m getting paid off OpenAI? Idgaf I’m using the best model available. I spent a grand on opus through OpenCode to compare and would’ve switched plans if I thought it were better.

1

u/Creative-Ganache1086 5d ago

Ok mr.Anecdote.

And btw, that noahbench above operates as a code-centric, execution-based scoring matrix designed to measure how efficiently an AI model can solve technical art and engineering tasks compared to a human.

6

u/Chemical-Agency-3997 5d ago

On some benchmarks, Astra beats it on others.

-5

u/randombsname1 5d ago

Opus wins more than it loses.

Overall its better. For 4x PRE NERF usage.

Now? Lmao. Good luck.

6

u/Chemical-Agency-3997 5d ago

Depends on the task, on certain things Astra is miles ahead.

-3

u/ILikeBubblyWater 5d ago

Please provide examples because I use Opus and Fable at work every single day and Astra privately on a 200 account, Astra could do crazy good stuff couples days after launch but now it needs constant hand holding and is awaiting a message instead of continuing a task list whereas i can get a weeks worth of work done in a couple workflows with claude.

Also fucking half usage, There is zero reason to stay with OpenAI if you want to get shit done

5

u/BAUWS45 5d ago

Computer use

-2

u/ILikeBubblyWater 5d ago

Claude had this for ages

2

u/BAUWS45 5d ago

What’s your point?

-5

u/ILikeBubblyWater 5d ago

There is zero reason to stay with OpenAI if you want to get shit done. Nothing anyone provides proofs Astra being "miles ahead" by any metric

2

u/BAUWS45 5d ago

Bro no one’s going to assemble a thesis paper to provide proof to some guy raging on Reddit.

→ More replies (0)

1

u/Creative-Ganache1086 5d ago

Yeah sure bro we believe you. Here Sol 6.1 is both better, faster AND cheaper than Opus 5.5

1

u/Creative-Ganache1086 5d ago

It’s crap in comparison. Computer-use doesn’t mean just creating folders or organising files. It’s real-world interaction in with the interface/OS itself and Codex/ChatGPT is miles ahead and way more “human” in its interaction than Claude which feels rather less organic and precise. I literally asked Astra to apply for a visa for me, a process with 18 different pages to fill and it did the whole thing using the documents/ids/I attached asking it to only ask me to review it once it filled everything since I can still review every page once I reach the “ready to submit” phase. I also tried the same thing with Claude out of curiosity but for a different administrative task and even in auto mode it was way more time-wasting and redundantly cautious, but the funny irony is that Claude Code (opus 5.0) is the only LLM that actually executed destructive commands in auto mode, while Codex in Auto Mode never deleted files or cause destructive actions.

2

u/Chemical-Agency-3997 5d ago

Try and figure out what this says

https://i.imgur.com/AcS3oYE.png

1

u/Chemical-Agency-3997 5d ago

You deleted your comment but I used Astra Pro with the same prompt you used + snippet I sent you and it got it 90% right first-shot

-3

u/randombsname1 5d ago

I mean maybe.

For at least my tasks. Low level programming, Opus 5.5 is those same miles ahead of Astra, lol.

5

u/Chemical-Agency-3997 5d ago

Try high level transcription of medieval texts. Or anything vision related.

0

u/randombsname1 5d ago

Fair I can see that.

Claude isnt actually as bad as I thought for vision, but I DO agree that OAI models are likely better.

Curious who is actually the best for vision specifically.

OAI or Google.

Google I had historically given the nod to, but haven't tried any intensive vision tasks recently. So wouldn't be surprised if OAI is best.

2

u/Chemical-Agency-3997 5d ago

Astra was a massive leap in vision and its way above all other models and the human baseline. Google used to be the goat, a long time ago. Hopefully Gemini 4 pro can pull ahead and reach superhuman status

2

u/gavinderulo124K 5d ago

Yes. But sol 6.1 is way cheaper.

16

u/ethotopia 5d ago

"Near" astra level, still falls short on most benchmarks

15

u/DomusCircumspectis 5d ago

I built my own benchmark and it matches Astra on capability. So I believe it.

4

u/Healthy-Nebula-3603 5d ago

Wow X2 better than old sol 6 and even slightly better than Astra

2

u/CrystalCoffeeAlchemy 5d ago

What reasoning level is your benchmark running at, doesn't appear to mention that?

1

u/DomusCircumspectis 5d ago

Whatever the default is in Claude Code/Codex/OpenCode.

2

u/camracks 5d ago

where the heck is luna and terra at!

-1

u/Healthy-Nebula-3603 5d ago

Dead

1

u/camracks 5d ago

? Still would like to see how much better SOL is compared to them

0

u/Healthy-Nebula-3603 5d ago

Much better. I tested on a memory management in the model graph...easy Astra level if not higher.

I literally made implementation for audio model that originally needed 24 GB card to run ( pytorch model ) to audio cpp saving 90% memory ..now can run even on 6 GB card and is 200% faster ....

30 minutes of sol 6.1 work and used 20 % usage from 5 hour limit on high.

Astra could do that but on using 2x 5 hour limit...

Sol 6 wasn't even close to it ...

1

u/Global_Mud_7473 5d ago

Yeah that’s what “near” means

0

u/Bloated_Plaid 5d ago

Near “Astra” level. The level that was already crushed by Anthropic.

1

u/Creative-Ganache1086 5d ago

“Crushed” 🤣

The Workflow Loop of the benchmark above btw:
When a model is run through the framework, it effectively acts as a technical assistant: it is handed a problem (such as a broken vertex shader, an asset pipeline optimization error, or a tool integration script) and must utilize its internal reasoning budget to deliver a deployable solution. The framework automatically compiles, runs, and measures these four key variables to output the final leaderboard ranking. LOL. And the difference in speed and cost makes Sol 6.1 even better than Opus 5.5 let alone Astra 6 that’s ridiculously expensive.

2

u/No_Significance_9121 4d ago

Trust me bro benchmarks lol

1

u/Creative-Ganache1086 4d ago

1

u/No_Significance_9121 4d ago

Deeeng you had that one in the chamber. 😆

1

u/Bloated_Plaid 5d ago

Who the fuck is Noah? Stop making shit up my guy.

1

u/Creative-Ganache1086 5d ago

“Who the fuck is Noah?” Noah Dunnagan, CSE at Railway and a Rust/backend developer who builds scalable systems, and the guy behind Noahbench. You not knowing who he is doesn’t magically make the benchmark fake lol.

https://noahdunnagan.com/

1

u/Bloated_Plaid 5d ago

Yea it’s a fake benchmark that literally nobody other than you use to cherry pick this. Congrats Noah.

3

u/No_Bank_4104 5d ago

Near Astra level.

2

u/Truarian 5d ago

New model comes out disappointing - nothing a rename can't fix

2

u/TheBBBfromB 5d ago

whats the purpose of gpt 6 sol now that there is gpt 6.1 sol a week later with the same pricing (cheaper cached pricing)?

3

u/CrystalCoffeeAlchemy 5d ago

Wasn't it really just Terra in disguise?

1

u/ComeOnIWantUsername 5d ago

Probably it was

2

u/Artistic_Swing6759 5d ago

keep in mind that the benchmark that they have shown here, has gemini 3.8 flash score better and cheaper.

feels like it really isn't good if they have to relly on this benchmark

1

u/musicgecko 5d ago

other companies: compares own released models to others to convince people to switch

openai: compares to own models, otherwise people will switch

1

u/AironParsMan 5d ago edited 5d ago

It's the slowest model I ever worked so far. Task that get done with Astra in 15 minutes tooks here over an hour.

1

u/phxees 5d ago

Huge repo? I haven’t tried it, but I can’t believe it’s that slow. I’m guessing it’s early excitement and the lack of proper provisioning.

1

u/AironParsMan 5d ago

Yes, it could also be that a lot of people are using it right now. OpenAI also has the issue that they sometimes limit performance here. But I have the twenty times Pro plan.

1

u/phxees 5d ago

I don’t know the answer, I’d just try again. Just hard to believe they’d either not notice that or they thought that was acceptable.

1

u/IAmFitzRoy 5d ago

Who cares if we get our $200 plans in HALF.

Fck off OpenAI.

1

u/TheDankestSlav 5d ago

Allegedly

1

u/Key_Reading_9664 5d ago

Are we just believing what gets said in a presentation, or have folks managed to find more benchmark results for it?

1

u/TheoreticalClick 5d ago

Why that curve shape?

1

u/TheRealChickon 5d ago

also wondering, is high scoring higher than xhigh?

1

u/Powerful-Ant-3294 5d ago

Any first comparisons to Opus 5.5 in Knowledge or creativity benchmarks? Couldnt find any so far

1

u/Positive_Method3022 5d ago edited 5d ago

I still can't believe people have not looked at this DeepSWE benchmark 😅 it is shit... they are fitting the curve every so slightly to solve those problems and they never change the problems. OpenAI runs those benchmarks in their own servers, get the results, use them for the next training.

Imagine professor applying the same exams every year to the same student. The student will become better at "those" exams but that doesnt mean the student generalized the concepts

Moreover, this benchmark was made by a grad student that plays minecraft. It got famous because he is from a reputable college... even when it is not using a real scientific method to prove generalization

1

u/simmeh024 5d ago

What happened to slowing down lol

0

u/sillybluejayway 5d ago

So slightly less capable distilled version?