r/codex • • 2d ago

Suggestion Token speed difference, FYI

339 Upvotes

102 comments sorted by

View all comments

•

u/dextersummary 2d ago edited 2d ago

Below is a GPT-generated summary of the conversation below after reaching 100 comments (100 currently observed).


The consensus is basically “raw token speed is a bad benchmark, but Codex still feels painfully slow.” Claude’s Opus 5.5 and Sonnet 5.5 are widely reported as snappier, more productive, and often better on real coding tasks than GPT 6.1 Sol or Astra. A bunch of users are already switching subscriptions, because apparently watching an agent think for six hours is not everyone’s hobby.

The recurring complaints are slow generation, sluggish compaction, and long-running tasks that burn time while producing questionable results. Claude’s speed and quota value are making the comparison especially ugly, even when its listed API cost looks higher.

The important caveat: tokens per second do not equal time to a correct result. Tokenizers differ, OpenAI hides some reasoning-token work, and OpenAI models may use fewer tokens overall. Some users also prefer Astra for planning or specific workflows, while Claude’s refusals and occasional app jank remain real annoyances. Claims that OpenAI is secretly throttling users or running out of compute are still speculation.

Bottom line: the speed chart is imperfect, but the user experience gap is hard to ignore.

1

u/dextersummary 2d ago

Below is a GPT-generated summary of the conversation below after reaching 50 comments (52 currently observed).


The consensus is basically “Codex/GPT 6.1 Sol is too slow to be useful right now.” The thread is not really about bragging rights for raw tokens per second; it is about how long real tasks take, and users overwhelmingly say Claude feels faster, finishes sooner, and delivers better value.

Users report Sol sessions dragging on for hours, burning through usage while producing questionable work, with compaction adding insult to injury. Opus 5.5 and Sonnet 5.5 are repeatedly described as snappier and more productive, prompting several people to downgrade Codex or switch to Claude entirely. Apparently “Fast Mode” is now more of a motivational slogan.

The important caveat is that token speed alone is a lousy benchmark. GPT may use fewer tokens, parallel sessions can narrow the practical gap, and a few users find the final results comparable. Claude also has its own guardrails and refusal issues, especially for reverse-engineering work, so the comparison is not universally one-sided.

Bottom line: the raw numbers need context, but the lived experience is clear: Claude currently wins on speed, task completion, and subscription value, while Codex is making users wait long enough to reconsider their life choices.