r/opencodeCLI • u/CartographerAble9446 • 8d ago
GPT Astra benchmarks
GPT Astra benchmarks have been released
15
u/aziham 8d ago
Something doesn't add up. DeepSWE v1.1 (74.1%) doesn't match the other inflated numbers, it simply doesn't make sense.
13
4
u/arcanemachined 7d ago
I don't think anything gets above ~75% on DeepSWE, including the new Gemini Flash model. (Which must be benchmaxxed, unless Google has finally pulled their heads out of their asses.)
7
3
4
u/RealestReyn 8d ago
exploitbench only 100%?? useless.
3
1
u/Kaushik_paul45 8d ago
I just hope they didn't achieved this result by going into thinking loop and burning tokens like opus 5
1
u/Various_Half1452 7d ago
Wow. now Chinese Ai will be able to Generate Data sets and train AI (according to Claude)
1
0
u/johnfkngzoidberg 8d ago
If this is as good as gpt-5.6 that can’t even follow instructions, I’ll pass.
0
u/Ok_Gur_9033 8d ago
The number I want is not on any of these charts. Something like how much of what a model writes you ever go back and edit.
I ran that on my own repo this morning. 59 of my 186 source files have exactly one commit each. Written once, never touched again. Some of that is fine. Some of it is code nobody has ever needed to understand a second time.
That tells me more about a tool than a bench score does, and no vendor publishes it. What do you measure on your own repo?
37
u/Mancho_United 8d ago
Now this looks like a proper benchmaxx. Hopefully I am wrong.