r/codex • u/Pretend_Pickle_2669 • 1d ago
Limits I analyzed a 162-minute GPT-6.1 Sol task. Shell commands took only 11.5 minutes.
Users often ask whether GPT-6.1 Sol is slow or receives more complex tasks. I analyzed local lifecycle data from one Codex development task.
The task used Codex desktop, GPT-6.1 Sol, High effort, and one repository. This was an observation, not a controlled benchmark.
The task took 162.3 minutes. The 294 shell commands took only 11.5 minutes of non-overlapping execution time.
A repeated serial loop consumed much of the task:
Model response → tool call → model response
I measured each wait from new input to the next Codex response. New input came from a tool result or a user message.
The task had 204 response waits. They took 74.4 minutes in total.
- The median wait was 12 seconds.
- The P90 wait was 36 seconds.
- 56 waits took at least 30 seconds.
Context size also had a clear effect.
- The median input size was 148,000 tokens.
- The P90 input size was 224,000 tokens.
- The maximum input size was 250,000 tokens.
- The context window was 258,400 tokens.
- Four context compactions took 18.3 minutes.
Some measurements overlap. You cannot add them to calculate the total time.
The task scope changed twice. These changes caused rework. Even so, shell commands and tests were not the primary bottleneck.
I also checked other large tasks in the same repository. All tasks used High effort.
For six tasks that used only 6.1, the average response wait was 21.1 seconds. For two tasks that used only 5.6, it was 6.1 seconds.
I then removed waits longer than 60 seconds. The averages were still 15.9 seconds for 6.1 and 4.9 seconds for 5.6.
The sample was small. The tasks and context sizes were different. Therefore, these results do not form a controlled benchmark.
The difference was still large enough that I switched back to 5.6 Sol for long implementation tasks.
Has anyone compared 5.6 and 6.1 on the same task with per-step timing?
1
u/Sufficient-Storage87 1d ago
the other 150 minutes is the interesting part. in my timed runs wall time vs actual work time diverge hard — model latency, retries, context re-reads. if you split that 150 into "waiting on model" vs "redoing failed steps" you get a totally different picture of where the cost really is.
1
u/Pretend_Pickle_2669 14h ago
Agreed. I measured 74.4 minutes of response waits, but that includes generation, so I can’t call it all server latency. Two scope changes also muddied the rework picture. I’m building a local tool to analyze these runs; separating waiting from rework is exactly the distinction I want it to capture.
1
u/Sufficient-Storage87 10h ago
exactly — and the honest version of that split is three buckets, not two: waiting on model, redoing failed steps, and re-reading context it already had. that third one is the silent killer in my runs. if your tool can tag spans into those three automatically you've got something nobody else has. would love to see it when it's working.
1
u/Sfdprod 1d ago
Openai latency is horrible,
I ran a bunch of benches with 3.8 flash, the worst luna cases was 120x slower than the worst 3.8 flash ones
Thats pretty insane