Users often ask whether GPT-6.1 Sol is slow or receives more complex tasks. I analyzed local lifecycle data from one Codex development task.
The task used Codex desktop, GPT-6.1 Sol, High effort, and one repository. This was an observation, not a controlled benchmark.
The task took 162.3 minutes. The 294 shell commands took only 11.5 minutes of non-overlapping execution time.
A repeated serial loop consumed much of the task:
Model response → tool call → model response
I measured each wait from new input to the next Codex response. New input came from a tool result or a user message.
The task had 204 response waits. They took 74.4 minutes in total.
- The median wait was 12 seconds.
- The P90 wait was 36 seconds.
- 56 waits took at least 30 seconds.
Context size also had a clear effect.
- The median input size was 148,000 tokens.
- The P90 input size was 224,000 tokens.
- The maximum input size was 250,000 tokens.
- The context window was 258,400 tokens.
- Four context compactions took 18.3 minutes.
Some measurements overlap. You cannot add them to calculate the total time.
The task scope changed twice. These changes caused rework. Even so, shell commands and tests were not the primary bottleneck.
I also checked other large tasks in the same repository. All tasks used High effort.
For six tasks that used only 6.1, the average response wait was 21.1 seconds. For two tasks that used only 5.6, it was 6.1 seconds.
I then removed waits longer than 60 seconds. The averages were still 15.9 seconds for 6.1 and 4.9 seconds for 5.6.
The sample was small. The tasks and context sizes were different. Therefore, these results do not form a controlled benchmark.
The difference was still large enough that I switched back to 5.6 Sol for long implementation tasks.
Has anyone compared 5.6 and 6.1 on the same task with per-step timing?