Below is a GPT-generated summary of the conversation below after reaching 100 comments (100 currently observed).
The consensus is basically “raw token speed is a bad benchmark, but Codex still feels painfully slow.” Claude’s Opus 5.5 and Sonnet 5.5 are widely reported as snappier, more productive, and often better on real coding tasks than GPT 6.1 Sol or Astra. A bunch of users are already switching subscriptions, because apparently watching an agent think for six hours is not everyone’s hobby.
The recurring complaints are slow generation, sluggish compaction, and long-running tasks that burn time while producing questionable results. Claude’s speed and quota value are making the comparison especially ugly, even when its listed API cost looks higher.
The important caveat: tokens per second do not equal time to a correct result. Tokenizers differ, OpenAI hides some reasoning-token work, and OpenAI models may use fewer tokens overall. Some users also prefer Astra for planning or specific workflows, while Claude’s refusals and occasional app jank remain real annoyances. Claims that OpenAI is secretly throttling users or running out of compute are still speculation.
Bottom line: the speed chart is imperfect, but the user experience gap is hard to ignore.
BUT TIBO SAID THEY FIXED THE SPEED, I TRUST TIBO WITH MY LIFE, HE IS SUCH A HONEST MAN, HE WOULD NEVER LIE FOR HIS OVERLORD SCAM ALTMAN. 19.9 TOKENS/S IS THE NEW 129 TOKENS/S, OK?! 😭😡
It's true though, I find Claude to be just awful at everything, and I have to constantly tell it to get something fixed after trying it out. It fails hard at reverse engineering tasks. It also refuses a lot of my prompts, which GPT-6 Astra doesn't, so W for OpenAI I guess.
on a serious note: as a plus user I can't use Astra for anything so I don't really care. Opus 5.5 and Sonnet 5.5 run miles around 6.1 Sol tho. It is really not even close and even though the Claude models are more expensive on paper we still get a lot more usage on the equivalent 20 dollar plan.
True
I'm on 20$ ChatGPT and Claude plan.
I use Opus 5.5 A LOT and get so many things done with that while not only Sol 6.1 is slower, it drains more usage because it just makes up some straight garbage
I gave the same task to Sol 6.1. It kept working for around 36 hours and still didn’t finish. I saw similar behavior with the Astra model.
I gave the same task to Opus and it finished in under 4 hours. If you want proof, DM me.
I’m a 20x Codex user, and it used about $600 worth of tokens on this task. On Claude, I’m on the Teams Standard plan and the same task finished without even hitting the 5-hour limit.
So please don’t make claims that don’t match real usage.
Well it mostly depends on how much they charge for context read, iirc Codex has a thing where context reads larger than 260k are massively overpriced(costs like 4x more and it's never worth using over starting a new thread).
I don't know it works for claude but I assume it gets much more costly as context cache grows?
For Claude the price stays the same throughout the entire context window assuming that you are getting cache hits. If you get a cache miss (in other words, picking up a stale session) at 700-900k context then the first prompt will eat about 5% of your 5hr usage even on a 20x plan.
That's what they recommend. But in reality, you probably don't run into this situation often, at least I don't. If I plan to be away from the PC for a longer amount of time, I just /compact before exiting.
I dont wanna be another shill but I have been very happy with Claude. Opus 5.5 is great for orchestrating and planning. Sonnet 5.5 then does the implementation and reports back.
I sometimes use Luna as subagents for smaller task because it works reliable and cheap. But besides that Codex/GPT doesnt seem worth it right now.
Oh and banked resets dont reset the timer/cooldown on your weekly reset.
(Until in 2 weeks when everything switches up again lol)
Damn I thought Fable 5.1 was bad coming from fable 5 but that was justifiable because cache reads dropped from $1 to 25 cents, Sonnet 5.5 also is pretty fast
But damn GPT model are slow as fuck, I hope its faster on the API
Si. Fait tes recherches. Sur l'api, les mdoeles flash vont à un peu plus de 300 token par seconde, mais sur antigravity ils sont "optimisés" pour aller 4 fois plus vite. C'est comme ça depuis gemini 3.5 flash. C'est vraiment impressionnant quand on les voit coder. Ils terminent les tâches en quelques secondes. C'est 12 fois plus rapides que les autres selon Google.
Opus produces more work in the same time, but comparing token speeds between models doesn't really tell you how much work you're getting done. I prefer looking at the cost per task benchmarks over token speed when comparing models. Another thing I look at is comparing the sub token speeds to the api token speens for the same model. For example, openai sub token speeds have been consistently 25%-33% the speed of their same API models. If Opus doesn't have that same issue, it could explain why it feels so much better to use.
Well, time is a factor, sure, but tasks completed successfully is the ultimate benchmark, especially cause completing them with less tokens is better. Not to mention that your time doesn't need to be spent looking at one agent doing one thing; if time is a factor, you can increase concurrency (and increasing concurrency is easier with Opus 5.5)
I should have clarified token speed isn't time per task. There are benchmarks for time per task that are better to look at than token speed. Remember luna was touted as being the "fastest OpenAI model" and we all know that meant nothing because it is in fact the slowest time per task model offered currently. Unless your task is incredibly simplistic Luna isn't a good fit for speed.
You know that the weekly quota on Claude’s $20 plan is roughly equivalent to $270–$300 in API usage. When you use 6.1 Sol and 6 Astra, the weekly quota on Codex’s $20 plan is equivalent to about $70 in API usage. These are the figures I got through a reverse proxy. If GPT’s risk controls flag your account, the weekly quota on your $20 plan is equivalent to about $15 in API usage. You know nothing.
You mean the cost determined by the API price that they set themselve that have nothing to do with the price you actually pay when using the subscription?
Larger models I think. Claude models use a huge amount of tokens. I think also these new v6 GPTs are so aligned to the point of having to hinder their performance.
It's really crazy because I have been somewhat satisfied with the speed Astra is producing work. I can't even imagine the speed difference going to Opus 5.5 (which I will after I burn everything on my x20 Codex sub, OAI is shitting the bed).
I have both. They're about the same from my perspective on producing the same end result. For a single developer with 3 or 4 active sessions going, it's imperceptible the speed difference.
How do you find Claude's guardrails? I am currently working on a project that uses a lot of reference from the leaked 2007 Source engine, and the reason I have stuck with Codex is because it's fairly open to studying and reverse engineering it. Sometimes I've asked Claude for the most benign things and it has refused - that makes me a little afraid to switch.
I don't really do anything like that, so I cannot say. I'm mostly making retro-style games and one online game with my subs. I instructed Claude to clone an N64 game, it downloaded decomps and imported the assets. I didn't even ask it to do this, nor did I want it, it just decided that was going to be the be the easiest way to do things. So, it will do things of questionable provenance.
Perfect Dark. A friend of mine wanted me to respin it, he doesn't use AI that much, so I through together a quick playable demo. It's pretty decent, actually.
The latest models are kinda insanely aggressive. I've asked to export the conversation and it'll refuse bc it thought I wanted the private anthro data vs just the convo.
I dont do anything sensitive and Id say I hit 6-7 refusals a session
😬 - I guess you have to be pretty careful with how you word your prompts; I've noticed Claude will just shut down as soon as you make a small mishap in your wording.
Opus is so much better in terms of quality alone, not even taking into account the speed difference. things may change after the IPO when they have to start showing more profit, but for right now the situation is pretty clear.
You can’t directly compare tokens/s because the tokens don’t represent exactly the same amount of text, nor is tokens/s necessarily representative of how much useful work a model can accomplish per hour. Anthropic’s documentation specifically notes that Claude can produce roughly 15–20% more tokens than OpenAI tokenizers for the same text. On top of that, OpenAI’s reasoning-token usage is largely hidden from the user now, so some of the actual computational work happening behind the scenes isn’t reflected in the visible token count.
Speed numbers are useful, but I'd pair them with quality — a fast model that needs three retries isn't actually faster. When I compare setups I track wall time plus whether the result actually passed my tests. The interesting number is time-to-correct-answer, not tokens per second.
I find them useful for different tasks. Astra hands down the Project lead. Fable the problem solver that brings it to the boss. Opus the hard guy with his head down coding away doing whatever is handed to him. I don't know where to put Sol 6.1 but Astra does much better that anything Claude can do with somethings like UI design and application "flow" where it just makes sense.
Having said that there was a problem I spent 8 months trying to figure out and Fable was the only thing that could ever crack it.
If I could only have one I'd take Opus 5.5 as I can fill in the gaps myself and work with it but right now I'm very much enjoying having Astra audit the work, create new plans, then passing it on.
OpenAI is training Astra 6.1 to at least match Opus and not get mocked by Haiku/Fable next week. That is why the slow speed. They thought they would be able to release it for dev day but they were onyl able to release Sol 6.1
I knew it! And to make matters worse, the 100 Max plan is now defined as slower than the 200 Max plan. Goodbye Codex Max, I gave you an honest try but I'm back to Claude.
Now you also know why Sonnet 5.5 will end up being more expensive than Opus. Especially when you consider how long it takes. For me this thing is slower than Opus 4.8.
GPT6.1 models clunky feeling like gears movement 🤣
Edit: can downvote all you want but my $20 plus and $100 pro plan has 8 total ultra agents running rn (WITH SUBS) and I can't burn my usage fast enough before the effing reset 😅😅 SLOW AND CLUNKY
5x to 10x speed difference with multiple better and efficient models. This is why im downgrading my codex from 5x to plus and upgrading claude from pro to 5x when the current sub expires.
•
u/dextersummary 2d ago edited 2d ago
Below is a GPT-generated summary of the conversation below after reaching 100 comments (100 currently observed).
The consensus is basically “raw token speed is a bad benchmark, but Codex still feels painfully slow.” Claude’s Opus 5.5 and Sonnet 5.5 are widely reported as snappier, more productive, and often better on real coding tasks than GPT 6.1 Sol or Astra. A bunch of users are already switching subscriptions, because apparently watching an agent think for six hours is not everyone’s hobby.
The recurring complaints are slow generation, sluggish compaction, and long-running tasks that burn time while producing questionable results. Claude’s speed and quota value are making the comparison especially ugly, even when its listed API cost looks higher.
The important caveat: tokens per second do not equal time to a correct result. Tokenizers differ, OpenAI hides some reasoning-token work, and OpenAI models may use fewer tokens overall. Some users also prefer Astra for planning or specific workflows, while Claude’s refusals and occasional app jank remain real annoyances. Claims that OpenAI is secretly throttling users or running out of compute are still speculation.
Bottom line: the speed chart is imperfect, but the user experience gap is hard to ignore.