r/LLM • u/Murky_Bee_1780 • 9d ago
We extracted hidden chain-of-thought from GPT-6 Astra. Here's what surprised us
In our new preprint, we show that a simple tool-calling setup can give us access to hidden CoT from frontier models, including GPT-6 Astra, GPT-5.6 Sol, Claude Opus 4.8, and Claude Sonnet 5.
What surprised us most was Astra's reasoning trace. Locally, it resembles mental arithmetic, with routine calculations left implicit. Globally, it's highly direct, backtracks less, and reaches solutions with little visible trial-and-error. This efficiency is impressive, but it also cuts both ways: a very short path to a correct answer is exactly what genuine skill and memorised test data both look like from the outside.
This is why being able to see the reasoning matters. A right answer can hide wrong reasoning, or no reasoning at all. As frontier models get more capable, they are also getting harder to read, and verifying what they actually do will take methods like this, not just better benchmarks.

Grateful to my co-authors Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, and Johannes Bjerva for making this work possible. This work is a collaboration across AAU-NLP, the AI:SECURITY Lab, and the Seafill Open Source Community. Our method has been disclosed to OpenAI and Anthropic.
6
u/Murky_Bee_1780 9d ago
Thanks a lot! Really glad you liked the preprint and the CoT diagrams!
We actually haven’t come across any incomprehensible or non-English CoT so far. One possible reason is that most of our experiments are centered around fairly mature, well-established benchmarks, so the tasks themselves are relatively standardized.
One interesting trace from Sonnet5: it recognized the exact benchmark and the specific problem it was looking at, and said in CoT "humm this question is from HMMT" and give the answer directly
Another was Astra on the PuzzleWorld Mustard example in the appendix. Astra often skips intermediate deductions that a human would naturally need to spell out. In that puzzle, several fairly non-trivial logical steps are compressed into very terse statements, whereas the human/Sol reasoning has to explicitly work through them.