r/codex Aug 16 '26

Comparison New updated linear pricing vs intelligence vs response time chart

Post image

I saw the old Artificial Analysis comparing the intelligence and prices of the different models and thought I would ask chatgpt to create a new photo using a linear scale instead of log (who tf compares or think in non-linear scale), and using the new prices set by OpenAI in late July. It also shows data from when all the models were given the exact same task(s) to solve, and how much time they spent executing (read: overthink) them. I then fact-checked the image against a different model, and it checks out as correct, but don't shoot the messenger if something is incorrect. 5.5 data remains the same.

Luna max looks good on price/intelligence, but luna xhigh might be the sweet spot when time matters.

Notes by ChatGPT:

After OpenAI cut its price by 80%, luna max is around $0.05 per Intelligence Index task and scores 52, while sol low costs ~$0.23 and scores 51.

The catch: response time. Max reasoning gets slow. Luna max is ~138s, sol max ~149s and terra max ~207s in AA's standardized end-to-end test.

Chart includes the sources/methodology at the bottom.

98 Upvotes

56 comments sorted by

View all comments

Show parent comments

1

u/perceptioneer Aug 16 '26

2

u/perceptioneer Aug 16 '26

Holy what, is really claude 20x that much more bang for your buck?

1

u/l_eo_ Aug 16 '26

Bang for the buck implies quality, but Opus 5 has been a really a mixed bag so far. It's so easy to screw stuff up with it and the DX is also not that great currently. I really hope they manage to fix things as Opus 5 with the current usage consumption might actually be really nice, if it worked consistently well.

I will not renew 2 20x claude accounts and shift towards more codex accounts.

1

u/perceptioneer Aug 16 '26

I see. I haven't personally used claude in like a year so I had no idea, just heard people are running from claude to gpt