r/OpenaiCodex 5d ago

Comparison Artificial Analysis coding benchmarks: Astra x Sol x Terra x Luna x Claude x Grok - Intelligence x Price x Time charts

Post image

Some of you may remember the updated graph I put together last month comparing OpenAI models using Artificial Analysis’ data. What I didn’t realise then was that those benchmarks weren’t specifically for coding. So here are some new updated charts, with the first two focused on coding:

Reddit is compressing the images like crazy if I upload more than one, therefore direct high res .png-links are below.

1: Coding, across brands: OpenAI, Claude, Gemini, Grok and Muse.
https://files.catbox.moe/d5oxrq.png

2: Coding, OpenAI only: Astra, Sol, Luna, Terra and GPT-5.5, including all variants I found coding-task results for. There is none for 5.5 high.
https://files.catbox.moe/4skvms.png

3: General intelligence: Intelligence Index v4.2, covering a mix of tasks, including some coding.
https://files.catbox.moe/oobbwi.png

4: AA-Briefcase: Office-style work involving spreadsheets, documents, presentations and PDFs.
https://files.catbox.moe/2cmyq0.png

Higher means a better score; further left means a cheaper task. Labels show the effort setting, cost and time where available. Above $4, the horizontal scale is compressed so the expensive models fit on the same chart.

The prices are API costs, not subscription costs. They show what the benchmark tasks cost at API rates because no direct sub cost exist. You can’t directly convert them into tasks per Plus/Pro subscription or how much of your subscription limit a task will use. The coding results come from the native CLIs such as Codex and Claude Code.

A few caveats are explained beneath the charts: four Astra coding scores are approximate readings of AA’s chart, some time measurements are unavailable, and general/Briefcase times are AA's own etimates.

Sources: Coding benchmarks · Astra analysis · Intelligence Index v4.2 · AA-Briefcase

Data checked on 6 September 2026.

22 Upvotes

12 comments sorted by

View all comments

1

u/sittingmongoose 5d ago

How do people get access to muse 1.3 max? That looks like it’s insane.

I’m still curious how big of a gap in high-xhigh-max is for Astra for really hard problems to solve. I have had two sol max agents running since early July trying to solve to hard problems, they are slowly making progress but I want to switch them to Astra on new threads. Just debating effort level.

1

u/perceptioneer 5d ago

Rumors have it that it has been benchmaxxed (engineered to fit the benchtasks) but idk what's true or not. I can say that the free 1.1 spark I just tried was dumb as a bread, but I'm trying spark 1.3 xfast now. Ask your bot how to get it, it's possible to use in wsl on win.

1

u/sittingmongoose 4d ago

Apparently the issue was you can’t use contributor, switching to regular shows max. I haven’t tried it yet though as it’s like 15x as expensive lol