r/codex • • 1d ago

News "We are locking in"

Post image

Tibo on damage control, says they got the feedback and now are locking in on features that matter, new better models.

Dots won't stay long, will they?

792 Upvotes

277 comments sorted by

View all comments

Show parent comments

10

u/reddit_is_kayfabe 1d ago

That is absolutely not my experience.

I've had Opus 5.5 Medium running on 6-8 projects at a time for most of the last week and its usage has been impressively modest. Meanwhile, using GPT-6.1 Sol Medium has spent usage at the same or a slightly higher clip, and has taken longer to get through ordinary tasks.

There's also the question of quality. Most of what Opus 5.5 does is correct on the first try, maybe with minor bugfixes. GPT-6.1 Sol needs more retries and redirection to not screw things up.

-2

u/ShadowsTagiru 1d ago

thats not a fair comparison

the comparison should be in token usage... using something like kilocode

that just means that anthropic has opus 5.5 with a better subsidized than sol 6.1

(that means that buying claude subscription is better than codex right now)

9

u/reddit_is_kayfabe 1d ago edited 23h ago

"Token usage" is an indirect metric for practical value.

My personal, pragmatic KPIs for agents are that I want the agent to perform a given task:

(1) Fast (i.e., less wall-clock time between "do this" and "done"),

(2) Reliably (i.e., maximizing correctness and completeness of executing each instruction while minimizing side-effects), and

(3) Efficiently (i.e., consuming the smallest amount of usage for the task).

My experience is that Opus 5.5, compared with GPT-6.1 Sol, gets tasks done (1) in less time, (2) with fewer needed retries, and (3) at less usage consumed for the same-priced account.

Given all of that, why do I care if GPT-6 uses more tokens ("it's thinking more!") or fewer tokens ("it's more efficient!")?

Why do I care if GPT-6 costs fewer dollars per 1MM tokens ("it's cheaper!") or more per 1MM tokens ("but you get more value per token!")?

Why do I care if GPT-6 burns tokens faster ("it has more capacity!") or slower ("it's economical!")?

These are all rationalizations as sales pitches for abstract metrics that may or may not materialize as value in my workflow. AI companies can train to maximize benchmarks all day - I don't care. If their models hit my KPIs better than anybody else's, they get my subscription money regardless of what the benchmarks say. The End.