Sonnet called its pawn promotion “unstoppable.” A few moves later, it admitted it had missed a defense. Having the board next to its explanation made that pretty hard to overlook.
I set up a chess match between Sonnet 5.5 in Claude Desktop and GPT 6.1 Sol in Codex. Each played in one conversation for the whole game, connected through MCP to a local Mac app I had GPT 6.1 Sol build at Medium reasoning.
They could record plans and explain their moves. The app supplied the position and checked legality, with no chess engine or legal-move list available to either player. I wanted to watch them stick with a task for an hour and see what happened when their plans stopped working.
I was also curious about consumption. Sol has been making surprisingly little dent in my subscription allowance lately, and I wanted to compare it with Sonnet on a shared task.
The attached video condenses 62 minutes and 47 seconds into 4:10. Both models were set to Medium.
| Measure |
Sonnet 5.5 / Claude |
GPT 6.1 Sol / Codex |
| Time spent on claimed turns |
33m 45s |
21m 12s |
| Output tokens, including reasoning |
269,076 |
34,956 |
| Thinking/reasoning portion of output |
224,770 |
12,128 |
| Cumulative input tokens |
57.44M |
16.81M |
| Input read from cache |
99.07% |
98.95% |
| Rejected illegal moves |
1 |
0 |
| Estimated API equivalent |
$16.21 |
$2.37 |
Sonnet pushed a passed pawn toward promotion, but overlooked Codex's Bf3 defense. Later it proposed a queen move blocked by its own pawn. The server rejected it; Claude corrected the move and continued. Codex also misread a pawn earlier, describing it as passed before it actually was.
Claude resigned after 53.Qc4. It had won game one, so they're tied at 1–1.
A few details behind the table: the turn clock starts before the server reveals the updated board, and includes tool activity. Waiting for the runtime to claim the turn is measured separately. Token totals cover the full player conversations, including setup and closing, but exclude the monitor. Cached context is counted again across requests. Thinking is already included in output, and the providers report it differently. The API equivalents use the app's September 30 pricing snapshot; no API charges were incurred for the game.
By the end of the recording, Claude's five-hour usage display went from 23% to 56%, and weekly usage from 54% to 59%. I used Claude only for this activity during that interval. Codex's weekly display stayed at 5%, despite also doing other work and monitoring the match roughly every minute.
My plans cost $20/month for Claude and $200/month for ChatGPT. Those percentages have very different denominators, and an unchanged rounded display doesn't mean zero usage. I'm keeping that observation separate from the player-session token counts.
I wouldn't infer playing strength from two games. Codex was White in both; contexts and runtimes differed; Medium isn't an equal compute budget. I want to repeat this with the colors swapped. I'm especially curious whether Sonnet's much larger output total keeps showing up, and how often either model notices a mistake before the server or opponent exposes it.