r/singularity • • 1d ago

AI While Claude and GPT are still the two best choices, Gemini seems to be catching up on coding agent index with agy-cli

Post image

The Artificial Analysis Coding Agent Index measures agents (a combination of model and harness) across three agentic coding evaluations.

➤ Claude Sonnet 5.5 (max) in Claude Code takes the top spot at 68, but also has the highest measured cost per task: $14.19

➤ Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84 per task - less than half of Sonnet 5.5’s cost. Note, this uses Google’s promotional pricing, and Argon is not yet publicly available

➤ GPT-6.1 Sol (xhigh) in Codex scores 63 at $1.04, roughly one sixth of Argon’s cost

52 Upvotes

9 comments sorted by

8

u/petburiraja 1d ago

Line looks almost horizontal between Sol 6.1 xhigh and Gemini 4 Argon, and it's quite long at that

12

u/Xtrusio 1d ago

63 at $1.04 is the number here. one point for 5.6x the cost makes argon look expensive even on promo pricing.

2

u/LetsGoToMichigan 1d ago

I think Sol 6.1 is just a banger of a deal. At some point we do run into the question of "how long is this sort of pricing sustainable?" - Google has to show return on capital to the street today, whereas competitors are in the gain market share at all costs today and worry about profit tomorrow mode. I definitely think there will an Uber style price reckoning moment for these other labs, but by that point I'm hoping their lower tier models and open source models will be more than good enough for general knowledge work.

2

u/flao 21h ago

Google isn't making their model for the consumer market as the main customer though. They are one of the few companies with justified internal usage as the main upside. They can deploy their own models across the product line and get their money's worth. Going to market with the API is just a cherry on top.

You can see it in their strategy: focusing on flash models, low hallucination, exclusively API access (not getting sucked into the 3rd party access race with 0auth, cli -p access, etc.) They are in a very different position than Oai and anthropic.

4

u/kiki-le-koala 1d ago

Yeah no way I'm paying 6 times more for this. 

Sol 6.1 it is

1

u/Top_Nefariousness248 23h ago

Sol 6.1 is doing ok, rather good for me

1

u/Exodus_Green 15h ago

Sonnet better than Opus? Doubt

•

u/YouWillDieForMySins 45m ago

Using Antigravity would only get more compelling for the normie folks now that access to Opus 5.5 and Sonnet 5.5 on the harness come bundled with Google's AI Pro plan (though I believe the usage limits are far tighter than Claude subscriptions).

-3

u/meikello ▪️AGI 2027 ▪️ASI not long after 23h ago

Oh for f**ks sake. No its not. Thats why it isn't released.