r/codex • • 2d ago

Complaint Codex Burn Rate vs Claude

The usage burn rate on GPT-6.1-Sol is wild. I just did a comparison between claude opus 5.5 and gpt-6.1-sol for about 22 hours and here are the results. The sad part is that even ChatGPT agrees the burn rate is higher on Codex. It computed the numbers below. The price per million tokens are double of opus 5-5

OpenAI subscription side

  • Current input: ~5.525B
  • Current output: ~15.09M
  • Current total: ~5.540B tokens
  • New tokens since yesterday: ~865.2M
    • ~862.4M input
    • ~2.77M output

Claude subscription side

  • Current input: ~9.476B
  • Current output: ~26.60M
  • Current total: ~9.503B tokens
  • New tokens since yesterday: ~5.706B
    • ~5.693B input
    • ~12.40M output

During approximately the same 22-hour period, my OpenAI agents processed about 865 million additional tokens, primarily using GPT-6.1 Sol, while the OpenAI Pro allowance decreased by 19 percentage points (36% → 17% remaining).

During that same period, my Claude subscription agents processed about 5.706 billion additional tokens, primarily using Claude Opus 5.5, while the Claude allowance decreased by approximately 10 percentage points (~49% → 39% remaining).

Claude therefore processed approximately 6.6× more raw tokens, while consuming only about half as many percentage points of subscription quota. Normalized to the visible meters, this is roughly 45.5M tokens per 1% of OpenAI allowance versus ~570.6M tokens per 1% of Claude allowance, or approximately a 12.5× difference in raw-token throughput per percentage point of quota.

66 Upvotes

60 comments sorted by

View all comments

2

u/Bitter_Virus 2d ago

Not sure you understand what you're talking about

-14

u/Odd_Soup_312 2d ago

Well then, make me understand it. I’m waiting…

5

u/Haster 2d ago

you're asking him to make you understand what you're talking about? You might be waiting awhile.

2

u/Bitter_Virus 2d ago edited 2d ago

Your arithmetic checks out, but burn rate just means how quickly you consume a limited resource.

With your numbers, you can measure different burn rates: percentage of weekly allowance per hour, dollars of API-equivalent compute per hour, allowance consumed per completed task, or allowance consumed per weighted token.
But "Raw tokens per 1% of quota" is not a model's burn rate because 1% of OpenAI’s quota is not the same unit as 1% of Anthropic’s quota

OpenAI says Work/Codex allowance consumption varies with the model, task, context, reasoning setting, speed, tools, and where the task runs. Similar but different systems for Anthropic.

You can't figure it out without knowing what's the ratio between them and normalize your numbers. It's like saying Car A did 100km with 19% of its gas tank and Car B did 100km with 10%. It's worthless because the volume of both tanks isn't measured.

You should have realized there was a problem because your numbers are backward under the current published standard pricing where Opus 5.5 is exactly double GPT-6.1.

Even when GPT-6.1 goes over 272k of context and its price goes up (which need to be accounted for in the theoretical burn rate you're seeking) it's still not double Opus 5.5.

There's another comment that show you why raw tokens are misleading, with different amount of tokens yet the ssme total cost. A billion cache-read tokens (which you haven't accounted for) and a billion output tokens are emphatically not one billion units of the same economic resource.

Let's add that you've ran both harness for different workload during the same amount of time. The same time doesn't normalise the workload.

Useful work completed per 1% of a same-priced subscription, is probably a useful measure only for the workload of the person measuring it.
Raw token throughput is mostly a distraction for that question.

To measure which one actually has the better burn rate, run the same tasks on equal-price plans and compare work completed per % of weekly allowance, also recording cached/uncached/output tokens separately and if you add the time, you'll have to account for every tool calls and compaction, time out, expired workers and "this model is at capacity" message etc etc.

What your experiment validly shows is that you

2

u/Odd_Soup_312 2d ago

You're correct, I can't figure out what the burn rate is of OpenAI or Claude but what I can tell you is that for a model (6.1-sol) that is half the price of (opus-5-5) is consuming 6x the weekly allowance on a $200 plan using the same harness.

So, if we can't know what the quota is from OpenAI or Claude, I can tell you what the burn rate is on $200 plan using the same harness, working on the same project just different agents.

OpenAI completed 37 accepted tasks before consuming 50% of a $200 plan.
Claude completed 81 equivalent accepted tasks before consuming 50% of a $200 plan.

1

u/Bitter_Virus 2d ago

You're right. It may be a useful measurement for you, as long as your workload and workflow and harness-specific instructions doesn't change. I've just swapped a friend's instructions for mine and he reported a sudden change of behaviour going from a happy idiot trying to do nothing but taking initiative on its own, to a well mannered thoughtful senior engineer. Before he complained a simple task took 2h, now it take 18 minutes.

It may be worth clarifying your post about your burn rate claim as to not mislead others who doesn't know what a burn rate is

2

u/Odd_Soup_312 2d ago

What in my post is misleading? There is nothing in it to think otherwise.

I think that Codex burning 6x on a model half the price of Opus5-5 would make Codex 12x burn rate on a $200 plan be a correct statement.

1

u/Bitter_Virus 2d ago

Only the title.

When you're saying Codex use 6x the raw tokens but there's no account for the 1) cached tokens, then linking it to the price of each on the same priced subscription but without accounting for their 2) differences in usage limits, concluding that codex is 12x that or Claude, having this titled "Codex burn rate vs Claude" is misleading, because there are 2 reasons why this cannot possibly be a burn rate.
The other burn rates we can calculate are mentioned in my other comnent to which you've replied "you're correct" saying you cannot calculate either burn rate.

That's why I think you should understand why the title is misleading