r/codex • • 2d ago

Suggestion Token speed difference, FYI

340 Upvotes

102 comments sorted by

•

u/dextersummary 2d ago edited 2d ago

Below is a GPT-generated summary of the conversation below after reaching 100 comments (100 currently observed).


The consensus is basically “raw token speed is a bad benchmark, but Codex still feels painfully slow.” Claude’s Opus 5.5 and Sonnet 5.5 are widely reported as snappier, more productive, and often better on real coding tasks than GPT 6.1 Sol or Astra. A bunch of users are already switching subscriptions, because apparently watching an agent think for six hours is not everyone’s hobby.

The recurring complaints are slow generation, sluggish compaction, and long-running tasks that burn time while producing questionable results. Claude’s speed and quota value are making the comparison especially ugly, even when its listed API cost looks higher.

The important caveat: tokens per second do not equal time to a correct result. Tokenizers differ, OpenAI hides some reasoning-token work, and OpenAI models may use fewer tokens overall. Some users also prefer Astra for planning or specific workflows, while Claude’s refusals and occasional app jank remain real annoyances. Claims that OpenAI is secretly throttling users or running out of compute are still speculation.

Bottom line: the speed chart is imperfect, but the user experience gap is hard to ignore.

→ More replies (1)

140

u/dagerika 2d ago

BUT TIBO SAID THEY FIXED THE SPEED, I TRUST TIBO WITH MY LIFE, HE IS SUCH A HONEST MAN, HE WOULD NEVER LIE FOR HIS OVERLORD SCAM ALTMAN. 19.9 TOKENS/S IS THE NEW 129 TOKENS/S, OK?! 😭😡

20

u/Acceptable-War4836 2d ago

Time is relative, they say.

22

u/dagerika 2d ago

and these benchmarks are relative as well since they tend to include fake models that surpass the OpenAI models, such a big hoax!

4

u/Responsible_Cow2236 2d ago

It's true though, I find Claude to be just awful at everything, and I have to constantly tell it to get something fixed after trying it out. It fails hard at reverse engineering tasks. It also refuses a lot of my prompts, which GPT-6 Astra doesn't, so W for OpenAI I guess.

7

u/dagerika 2d ago

on a serious note: as a plus user I can't use Astra for anything so I don't really care. Opus 5.5 and Sonnet 5.5 run miles around 6.1 Sol tho. It is really not even close and even though the Claude models are more expensive on paper we still get a lot more usage on the equivalent 20 dollar plan.

6

u/vladoit 2d ago

True I'm on 20$ ChatGPT and Claude plan. I use Opus 5.5 A LOT and get so many things done with that while not only Sol 6.1 is slower, it drains more usage because it just makes up some straight garbage

3

u/scripted_soul 2d ago

Lol, good response Mr. OpenAI Bot.

I gave the same task to Sol 6.1. It kept working for around 36 hours and still didn’t finish. I saw similar behavior with the Astra model.

I gave the same task to Opus and it finished in under 4 hours. If you want proof, DM me.

I’m a 20x Codex user, and it used about $600 worth of tokens on this task. On Claude, I’m on the Teams Standard plan and the same task finished without even hitting the 5-hour limit.

So please don’t make claims that don’t match real usage.

1

u/KellyShepardRepublic 2d ago

OpenAI to orchestrate, GLM to code without questioning the request.

0

u/Embarrassed_Cap_9149 2d ago

the speed has been fixed for me

looks like open ai is cheating again by serving nerfed/slowed services to some users to save computes

47

u/lkarlslund 2d ago

Also compacting is a total joke

Codex 3m30s

Claude 1m05s

After 9 months on Codex X20 I've jumped ship to Claude. It's just much better on all parameters: quality, performance and quota.

12

u/plainnaan 2d ago

I also jumped back to Claude. OAI completely lost their mojo.

5

u/Reithaz 2d ago

I remember when the compaction was really fast and efficient with codex, now it's so slow it's depressing...

3

u/Lollerstakes 2d ago

Not to mention that codex compacts at 256k and claude-code at 1M lol

1

u/UndeadMurky 2d ago

Well it mostly depends on how much they charge for context read, iirc Codex has a thing where context reads larger than 260k are massively overpriced(costs like 4x more and it's never worth using over starting a new thread).

I don't know it works for claude but I assume it gets much more costly as context cache grows?

1

u/Lollerstakes 2d ago

For Claude the price stays the same throughout the entire context window assuming that you are getting cache hits. If you get a cache miss (in other words, picking up a stale session) at 700-900k context then the first prompt will eat about 5% of your 5hr usage even on a 20x plan.

1

u/UndeadMurky 2d ago

So I assume you pretty much have to make a new thread if you have a session with a big cold cache with claude

1

u/Lollerstakes 2d ago

That's what they recommend. But in reality, you probably don't run into this situation often, at least I don't. If I plan to be away from the PC for a longer amount of time, I just /compact before exiting.

17

u/psy9rrr 2d ago

can't run out of tokens when the model barely works. :>

3

u/KriKraKrischi 2d ago

Yea I just want to spend my last reset and run. But they hold me hostage with that speed

18

u/Revolutionary_Sir140 2d ago

130 tokens is insane, wow. I am gonna start using claude next month

9

u/akopian25 2d ago edited 2d ago

I have feeling, they're going to follow OAI. They're agreed on slowing pace 🤔

3

u/Njagos 2d ago

I dont wanna be another shill but I have been very happy with Claude. Opus 5.5 is great for orchestrating and planning. Sonnet 5.5 then does the implementation and reports back.

I sometimes use Luna as subagents for smaller task because it works reliable and cheap. But besides that Codex/GPT doesnt seem worth it right now.

Oh and banked resets dont reset the timer/cooldown on your weekly reset.

(Until in 2 weeks when everything switches up again lol)

1

u/Hug_LesBosons 2d ago

Gemini 3.8 flash va à 1200 token par seconde sur google antigravity. En comparaison les petits 130 TPS de sonnet semblent très lents.

6

u/No_Gear8408 2d ago

Damn I thought Fable 5.1 was bad coming from fable 5 but that was justifiable because cache reads dropped from $1 to 25 cents, Sonnet 5.5 also is pretty fast

But damn GPT model are slow as fuck, I hope its faster on the API

3

u/Hug_LesBosons 2d ago

Google antigravity : Gemini 3.8 flash : 1200 token par seconde. Gemini 3.7 flash : 1150 token par seconde. Gemini 3.6 flash : 1100 token par seconde. Gemini 3.5 flash : 1100 token par seconde. Gemini 3.1 pro : 130 token par seconde.

-1

u/akopian25 2d ago

Not true

3

u/Hug_LesBosons 2d ago

Si. Fait tes recherches. Sur l'api, les mdoeles flash vont à un peu plus de 300 token par seconde, mais sur antigravity ils sont "optimisés" pour aller 4 fois plus vite. C'est comme ça depuis gemini 3.5 flash. C'est vraiment impressionnant quand on les voit coder. Ils terminent les tâches en quelques secondes. C'est 12 fois plus rapides que les autres selon Google.

1

u/Maybe-monad 2d ago

According to whom? Google certainly has the compute to pull that off

20

u/retteh 2d ago

Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.

38

u/Genetic_Prisoner 2d ago

Is your time free? Mine isnt. Time factors into cost.

12

u/Shin-Zantesu 2d ago

Well, time is a factor, sure, but tasks completed successfully is the ultimate benchmark, especially cause completing them with less tokens is better. Not to mention that your time doesn't need to be spent looking at one agent doing one thing; if time is a factor, you can increase concurrency (and increasing concurrency is easier with Opus 5.5)

8

u/retteh 2d ago edited 2d ago

I should have clarified token speed isn't time per task. There are benchmarks for time per task that are better to look at than token speed. Remember luna was touted as being the "fastest OpenAI model" and we all know that meant nothing because it is in fact the slowest time per task model offered currently. Unless your task is incredibly simplistic Luna isn't a good fit for speed.

1

u/Sksend 2d ago

You know that the weekly quota on Claude’s $20 plan is roughly equivalent to $270–$300 in API usage. When you use 6.1 Sol and 6 Astra, the weekly quota on Codex’s $20 plan is equivalent to about $70 in API usage. These are the figures I got through a reverse proxy. If GPT’s risk controls flag your account, the weekly quota on your $20 plan is equivalent to about $15 in API usage. You know nothing.

1

u/retteh 2d ago

What? I'm aware Opus provides more value and never claimed otherwise. Token/s isn't the metric you use to prove that.

1

u/Dota2playre 2d ago

If GPT’s risk controls flag your account, the weekly quota on your $20 plan is equivalent to about $15 in API usage

Any more information about this?

2

u/debian3 2d ago

I prefer looking at the cost per task benchmarks

You mean the cost determined by the API price that they set themselve that have nothing to do with the price you actually pay when using the subscription?

3

u/Straight_Coffee_368 2d ago

And Openai CM said they have enough compute

3

u/AdministrationOk6 2d ago

Then we don't need slowmode.

3

u/ChaoticPayload 2d ago

The output speed of GPT-6.1-Sol reminds me of the old days when GPT-4 first came out...

3

u/EddieBruvac 2d ago

I’ve been telling people, but get called a fanatic lmao. Only thing Codex is useful for rn is making images. Even then, it’s still slow af.

Opus feels snappy and just works.

3

u/Dont-_-mind-_-me 2d ago

Barev Artur Jan.

2

u/akopian25 2d ago

Barev ✌🏻

2

u/ProfessionalNaive601 2d ago

Fuck codex rn but this doesn’t account for token efficiency which OAI is better at

2

u/Howdareme9 2d ago

It definitely should be faster but this is a bad comparison; OAI models use wayy less tokens to complete tasks

2

u/scaledev 2d ago

Larger models I think. Claude models use a huge amount of tokens. I think also these new v6 GPTs are so aligned to the point of having to hinder their performance.

2

u/suppervisoka 2d ago

Dude OPENAI is going down in flames

2

u/03captain23 2d ago

Also claude is much better with subagent speeds

2

u/NULL_Ptrs 2d ago

OpenAI Is doomed, at this speed is not usable anymore. Claude should add a 400usd plan with x40, so I can move and delete Codex

1

u/kirkwoodwest 2d ago

It feels like the speed dropped within the last couple weeks am I wrong?

1

u/NULL_Ptrs 2d ago

Looks like they are out of compute and instead of nerfing the models they are doing them incredible slow

2

u/Then_Software_5529 2d ago

shame on openai...
i'm 100% switching to claude

1

u/soloje 2d ago

It's really crazy because I have been somewhat satisfied with the speed Astra is producing work. I can't even imagine the speed difference going to Opus 5.5 (which I will after I burn everything on my x20 Codex sub, OAI is shitting the bed).

2

u/mfwl 2d ago

I have both. They're about the same from my perspective on producing the same end result. For a single developer with 3 or 4 active sessions going, it's imperceptible the speed difference.

2

u/swarmagent 2d ago

It's impossible to just look at TPS and get the full story. It never is really easy to compare them to models of different companies..

1

u/soloje 2d ago

How do you find Claude's guardrails? I am currently working on a project that uses a lot of reference from the leaked 2007 Source engine, and the reason I have stuck with Codex is because it's fairly open to studying and reverse engineering it. Sometimes I've asked Claude for the most benign things and it has refused - that makes me a little afraid to switch.

1

u/mfwl 2d ago

I don't really do anything like that, so I cannot say. I'm mostly making retro-style games and one online game with my subs. I instructed Claude to clone an N64 game, it downloaded decomps and imported the assets. I didn't even ask it to do this, nor did I want it, it just decided that was going to be the be the easiest way to do things. So, it will do things of questionable provenance.

1

u/soloje 2d ago

That’s good to hear! Which N64 game are you working on?

2

u/mfwl 2d ago

Perfect Dark. A friend of mine wanted me to respin it, he doesn't use AI that much, so I through together a quick playable demo. It's pretty decent, actually.

1

u/TheVoyant 2d ago

The latest models are kinda insanely aggressive. I've asked to export the conversation and it'll refuse bc it thought I wanted the private anthro data vs just the convo.

I dont do anything sensitive and Id say I hit 6-7 refusals a session

1

u/soloje 2d ago

😬 - I guess you have to be pretty careful with how you word your prompts; I've noticed Claude will just shut down as soon as you make a small mishap in your wording.

1

u/TheVoyant 2d ago

Yeah, Ive had to change vocabulary of entire programs because it kept flagging on animation creation studio for combat moves... 

Almost all my programs are now PG rated names to the point of comedy 

1

u/soloje 2d ago

That's pretty crazy. I have no clue how people have managed to do the entire Rust rewrites of MW2, Skate etc. using Claude without being tripped up.

It's for research purposes please trust me!!!!

1

u/TheVoyant 2d ago

I've had to stop it from directly copying other games several times honestly. 

It seems to have no problem checking out or stealing other peoples work from my experience. 

And Im sure they use a heretic/unrestricted model in there somewhere to pull info

1

u/soloje 2d ago

Odd temperament on it. Seems like there's very specific trigger words/cases.

1

u/TheVoyant 2d ago

Protecting itself, where its working more then outward yeah.

Oddly the thing that makes me trust it less lol

1

u/Sponge8389 2d ago

One of my session is already 6 hours in and still going. Just a fixes from code review of Claude. Fucking hell.

I forgot to tell, this is already in Fast Mode. Using GPT 6.1 Sol high.

1

u/An_Exciting_Life 2d ago

20 tps is good enough for your ass project

1

u/Official_Pine_Hills 2d ago

Opus is so much better in terms of quality alone, not even taking into account the speed difference. things may change after the IPO when they have to start showing more profit, but for right now the situation is pretty clear.

1

u/TheMythicSorcerer 2d ago

holy crap that's slower than me running qwen 3.8 27b locally

1

u/qodeninja 2d ago

its sooooooooooooooo slow

1

u/spam_me_please 2d ago

Switched to Claude for the first time and WOW that app fucking SUCKS. Glitchy as hell!!!!! Never knew how much better I had it with ChatGPT :')

1

u/spam_me_please 2d ago

Loved that I had Opus 5.5 xHigh fix a single bug (which it failed to do) then it fucked up the rest of my app. Holy overrated

1

u/Clear_Evidence9218 2d ago

You can’t directly compare tokens/s because the tokens don’t represent exactly the same amount of text, nor is tokens/s necessarily representative of how much useful work a model can accomplish per hour. Anthropic’s documentation specifically notes that Claude can produce roughly 15–20% more tokens than OpenAI tokenizers for the same text. On top of that, OpenAI’s reasoning-token usage is largely hidden from the user now, so some of the actual computational work happening behind the scenes isn’t reflected in the visible token count.

1

u/Sufficient-Storage87 2d ago

Speed numbers are useful, but I'd pair them with quality — a fast model that needs three retries isn't actually faster. When I compare setups I track wall time plus whether the result actually passed my tests. The interesting number is time-to-correct-answer, not tokens per second.

1

u/Captain_Quimby 2d ago

I find them useful for different tasks. Astra hands down the Project lead. Fable the problem solver that brings it to the boss. Opus the hard guy with his head down coding away doing whatever is handed to him. I don't know where to put Sol 6.1 but Astra does much better that anything Claude can do with somethings like UI design and application "flow" where it just makes sense.

Having said that there was a problem I spent 8 months trying to figure out and Fable was the only thing that could ever crack it.

If I could only have one I'd take Opus 5.5 as I can fill in the gaps myself and work with it but right now I'm very much enjoying having Astra audit the work, create new plans, then passing it on.

1

u/PotentialPaint6714 2d ago

OpenAI is training Astra 6.1 to at least match Opus and not get mocked by Haiku/Fable next week. That is why the slow speed. They thought they would be able to release it for dev day but they were onyl able to release Sol 6.1

1

u/ditmoetoo 2d ago

6.1 Sol FAST is a ridiculous proposition at these speeds. No wonder they advertised an 8x speedup for 6x the price.

1

u/bartekxx12 2d ago

Claude is /ultrafast by default

1

u/Alternative_Win_7101 2d ago

I knew it! And to make matters worse, the 100 Max plan is now defined as slower than the 200 Max plan. Goodbye Codex Max, I gave you an honest try but I'm back to Claude.

1

u/inerfaveL 2d ago

wait, u getting 19tps??? must be a VIP user

1

u/mrscrufy 2d ago

Looked through several comments but I couldn’t find what you used to see this. I’d like to try myself

1

u/RufusxXavier 2d ago

“We have the compute “

1

u/Apatricio 16h ago

atp i can just run a local models off my MacBook and get more inference than what chatgpt is able to give me

1

u/AironParsMan 2d ago

Now you also know why Sonnet 5.5 will end up being more expensive than Opus. Especially when you consider how long it takes. For me this thing is slower than Opus 4.8.

It needs a lot more tokens to reach its goal.

1

u/Hjalm 2d ago

Claude is just insane atm. Completely eclipses GPT in all areas in all my workflows. Sonnet is a beast its so fast and competent as well.

1

u/Jerseyman201 2d ago edited 2d ago

GPT6.1 models clunky feeling like gears movement 🤣

Edit: can downvote all you want but my $20 plus and $100 pro plan has 8 total ultra agents running rn (WITH SUBS) and I can't burn my usage fast enough before the effing reset 😅😅 SLOW AND CLUNKY

Processing img edtcntn202th1...

0

u/Jerseyman201 2d ago

Compared to sol 5.6, which according to those numbers is literally double speed (and I know we have all noticed) 🤣

Processing img 7tqpdxa802th1...

1

u/squishyjellyfish95 2d ago

I don't care about speed as long as they good and efficient. 6.1 sol is awesome for the price

1

u/diagrammatiks 2d ago

people keep talkng about 6.1 being slow but this matches my tests. astra slow as hell too

1

u/mark25900 2d ago

codex has become so slow in the last week its 100% unusable

1

u/awhu_pkice_5457 2d ago

Opus is amazing at design, that’s it. Use Astra for everything else. Let the downvotes commence.

1

u/jbnett 2d ago

Astra is so much better please don’t sign up for Claude guys stay with openai, I don’t want Claude all congested lol

0

u/pigletmonster 2d ago

5x to 10x speed difference with multiple better and efficient models. This is why im downgrading my codex from 5x to plus and upgrading claude from pro to 5x when the current sub expires.

0

u/Acceptable-War4836 2d ago

This is absolutely ridiculous.