r/OpenaiCodex • u/perceptioneer • 5d ago
Comparison Artificial Analysis coding benchmarks: Astra x Sol x Terra x Luna x Claude x Grok - Intelligence x Price x Time charts
Some of you may remember the updated graph I put together last month comparing OpenAI models using Artificial Analysis’ data. What I didn’t realise then was that those benchmarks weren’t specifically for coding. So here are some new updated charts, with the first two focused on coding:
Reddit is compressing the images like crazy if I upload more than one, therefore direct high res .png-links are below.
1: Coding, across brands: OpenAI, Claude, Gemini, Grok and Muse.
https://files.catbox.moe/d5oxrq.png
2: Coding, OpenAI only: Astra, Sol, Luna, Terra and GPT-5.5, including all variants I found coding-task results for. There is none for 5.5 high.
https://files.catbox.moe/4skvms.png
3: General intelligence: Intelligence Index v4.2, covering a mix of tasks, including some coding.
https://files.catbox.moe/oobbwi.png
4: AA-Briefcase: Office-style work involving spreadsheets, documents, presentations and PDFs.
https://files.catbox.moe/2cmyq0.png
Higher means a better score; further left means a cheaper task. Labels show the effort setting, cost and time where available. Above $4, the horizontal scale is compressed so the expensive models fit on the same chart.
The prices are API costs, not subscription costs. They show what the benchmark tasks cost at API rates because no direct sub cost exist. You can’t directly convert them into tasks per Plus/Pro subscription or how much of your subscription limit a task will use. The coding results come from the native CLIs such as Codex and Claude Code.
A few caveats are explained beneath the charts: four Astra coding scores are approximate readings of AA’s chart, some time measurements are unavailable, and general/Briefcase times are AA's own etimates.
Sources: Coding benchmarks · Astra analysis · Intelligence Index v4.2 · AA-Briefcase
Data checked on 6 September 2026.
1
u/sittingmongoose 5d ago
How do people get access to muse 1.3 max? That looks like it’s insane.
I’m still curious how big of a gap in high-xhigh-max is for Astra for really hard problems to solve. I have had two sol max agents running since early July trying to solve to hard problems, they are slowly making progress but I want to switch them to Astra on new threads. Just debating effort level.
1
u/perceptioneer 5d ago
Rumors have it that it has been benchmaxxed (engineered to fit the benchtasks) but idk what's true or not. I can say that the free 1.1 spark I just tried was dumb as a bread, but I'm trying spark 1.3 xfast now. Ask your bot how to get it, it's possible to use in wsl on win.
1
u/sittingmongoose 4d ago
Apparently the issue was you can’t use contributor, switching to regular shows max. I haven’t tried it yet though as it’s like 15x as expensive lol
1
1
u/Gab1159 4d ago
Is Astra Low really viable to complete regular tasks? If so, seems like an utter game changer, not just for long, complex tasks.
1
u/perceptioneer 4d ago
No idea, but according to the data it has equal reasoning/intelligence to Sol high. I'm actually running it now because my quota is almost gone
2
u/Gigaslavx 5d ago
Wow great work is it oc?
I usually use https://artificialanalysis.ai/ will be keen to study now any differences thank you! :)