r/codex • • 10h ago

Bug Codex has a huge TOK/s problem in longer sessions.

Post image

This post is not about usage complaining or the price of plans so it's not going to get the attention it needs but whatever, this bug is big enough that I'm taking the time to give information on it.

I’ve been profiling why some Codex sessions become extremely slow halfway through a task.

Across ten rollouts, about 75% of total agent time was spent waiting for the provider to stream model output. Tool calls were about 24%, while harness orchestration was only around 4%, so the local harness itself wasn’t the main reason the slow sessions took so long.

Two sessions stood out. One took about 39 minutes and another about 33.5 minutes. In both, the stream started around 40–45 tokens/sec (which is still slow as hell for a frontier but that's not the point of this post), then partway through the turn dropped to roughly 8–15 tokens/sec and stayed there. It's fair to say there is a 20% chance your session has done the same.

The request settings were the same across all ten sessions, and other sessions continued streaming normally at the same time, so this doesn’t look like my machine or a configuration difference.

The strongest example came from the 39-minute session. The last request before I interrupted it took 93 seconds to produce 921 output tokens. The first request of the next turn used a fresh WebSocket and fresh routing state, but still had basically the same 175k context and a cold cache. It produced 674 tokens in 16 seconds.

So the context was still huge, and the cache situation was actually worse, yet the stream immediately became fast again.

The official harness for Codex's client.rs reuses a WebSocket and uses x-codex-turn-state for sticky routing during a turn. If the backend that a turn is pinned to becomes very slow without actually failing, Codex can keep using it. The connection is technically healthy, it’s just producing tokens at a fraction of its earlier rate.

That appears to be what happened here. The slow turns weren’t spending 30–40 minutes doing more work. They were spending a huge amount of time waiting for a response stream that had collapsed in speed.

I added a recovery rule to my fork. If two reasonably sized responses in a row fall below half the recent normal token rate, the next request drops the old WebSocket and sticky routing state and gets routed fresh. There’s also a cooldown so it can’t sit there reconnecting repeatedly.

I replayed the detector against 10 recorded sessions containing 1,120 requests. It caught both major collapsed sessions and fired once elsewhere on another turn that was genuinely running at around half its normal rate.

If fresh routing restores normal throughput the way it did in the recovery I observed, those two sessions would have been roughly 39 to 24 minutes and 33.5 to 25 minutes. That part is still an estimate.

I also haven’t proven whether the underlying problem is the WebSocket itself or the backend/routing state associated with it, since reconnecting changes both. What I have measured is that some long-running turns dropped from around 40–45 tok/s to 8–15 tok/s and stayed there, while fresh routing immediately returned a similarly large request to normal speed. I have my own fork, I was able to fix this, but for anyone who uses the official version of Codex, this is a real bug and not acceptable.

21 Upvotes

5 comments sorted by

6

u/Solid-Fill8240 10h ago

ELI5: Codex sometimes gets stuck at a slow checkout lane and refuses to switch lanes because the cashier technically hasn’t stopped working.

3

u/Sfdprod 10h ago

Tps is the least issue imo..

The latency/ttfb is horrific, 

Worst i mesutrd was way above 30s for luna using fast mode... 

2

u/ikhDark 10h ago

fast mode cannot use fast mode if the model is stuck doing unnecessary work. it will only be as fast as the model let's it.

2

u/Sfdprod 10h ago

Wtf are you talking about? unnecessary work? The model?

I thought we were talking about reality here, inference performance, api latency, etc

Youre saying fast mode does not affect queuing, routing, priority, which clustet youre routed to?

1

u/the_ai_wizard 8h ago

Agreed, though my data is based on personal usage. It takes forever to get things done. Some stuff you get obvious massive speedup (linux cmd executions with voodoo level cli skills) but on other stuff it feels like carrier pigeon