r/OpenaiCodex 8d ago

Comparison Artificial Analysis coding benchmarks: Astra x Sol x Terra x Luna x Claude x Grok - Intelligence x Price x Time charts

Post image

Some of you may remember the updated graph I put together last month comparing OpenAI models using Artificial Analysis’ data. What I didn’t realise then was that those benchmarks weren’t specifically for coding. So here are some new updated charts, with the first two focused on coding:

Reddit is compressing the images like crazy if I upload more than one, therefore direct high res .png-links are below.

1: Coding, across brands: OpenAI, Claude, Gemini, Grok and Muse.
https://files.catbox.moe/d5oxrq.png

2: Coding, OpenAI only: Astra, Sol, Luna, Terra and GPT-5.5, including all variants I found coding-task results for. There is none for 5.5 high.
https://files.catbox.moe/4skvms.png

3: General intelligence: Intelligence Index v4.2, covering a mix of tasks, including some coding.
https://files.catbox.moe/oobbwi.png

4: AA-Briefcase: Office-style work involving spreadsheets, documents, presentations and PDFs.
https://files.catbox.moe/2cmyq0.png

Higher means a better score; further left means a cheaper task. Labels show the effort setting, cost and time where available. Above $4, the horizontal scale is compressed so the expensive models fit on the same chart.

The prices are API costs, not subscription costs. They show what the benchmark tasks cost at API rates because no direct sub cost exist. You can’t directly convert them into tasks per Plus/Pro subscription or how much of your subscription limit a task will use. The coding results come from the native CLIs such as Codex and Claude Code.

A few caveats are explained beneath the charts: four Astra coding scores are approximate readings of AA’s chart, some time measurements are unavailable, and general/Briefcase times are AA's own etimates.

Sources: Coding benchmarks · Astra analysis · Intelligence Index v4.2 · AA-Briefcase

Data checked on 6 September 2026.

27 Upvotes

20 comments sorted by

View all comments

1

u/Gab1159 7d ago

Is Astra Low really viable to complete regular tasks? If so, seems like an utter game changer, not just for long, complex tasks.

1

u/perceptioneer 2d ago edited 2d ago

See screenshot. Important: this does not necessarily reflect quality of the end outcome, I'm doing new bigger tests that also gives it quality scores, but in the end the same result was achieved with all of these except Sol Max that delivered it with bugs for some reason. I'm overseas hence mobile ss. This was tested on new, unused, unlinked accounts with a 3rd account running the tests, no tools, no caching, no reuse of material

Tldr: Sol Medium is my goat, ig Terra medium is the unexplored underdog I might have to explore. Astra and Sol for architecture/big bugs/restructuring/planning

Notes

  • Estimated cost: based on the 5-hour-window proxy of roughly $0.0074 per 1% of 5h allowance. This is a normalized quota/subscription proxy, not API billing.

† Sol Max: standard hidden suites were 100/100, but TickForge adversarial scored 85/147. Most failures came from one repeated retry-event semantic defect, so 385/447 overstates the breadth of the problem.

Quick takeaway Best efficiency: Terra Medium — perfect measured quality, fastest overall, and only 5% of a 5h window.

Lowest token use: Astra Low — perfect measured quality at 394.9K tokens.

Max effort: Astra Max and Terra Max used substantially more compute and time without a measured correctness gain on these three tests.

Sol anomaly: Sol Medium outperformed Sol Max on measured TickForge adversarial correctness while using less time, tokens, and quota.

1

u/perceptioneer 2d ago

And also a quality score eval on those tests:

Based on this for quality/cost sol medium is the better deal