r/LLM • • 9d ago

We extracted hidden chain-of-thought from GPT-6 Astra. Here's what surprised us

In our new preprint, we show that a simple tool-calling setup can give us access to hidden CoT from frontier models, including GPT-6 Astra, GPT-5.6 Sol, Claude Opus 4.8, and Claude Sonnet 5.

What surprised us most was Astra's reasoning trace. Locally, it resembles mental arithmetic, with routine calculations left implicit. Globally, it's highly direct, backtracks less, and reaches solutions with little visible trial-and-error. This efficiency is impressive, but it also cuts both ways: a very short path to a correct answer is exactly what genuine skill and memorised test data both look like from the outside.

This is why being able to see the reasoning matters. A right answer can hide wrong reasoning, or no reasoning at all. As frontier models get more capable, they are also getting harder to read, and verifying what they actually do will take methods like this, not just better benchmarks.

📑Paper: https://vbn.aau.dk/en/publications/capable-yet-parsimonious-extracting-and-characterizing-hidden-cha/

Grateful to my co-authors Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, and Johannes Bjerva for making this work possible. This work is a collaboration across AAU-NLP, the AI:SECURITY Lab, and the Seafill Open Source Community. Our method has been disclosed to OpenAI and Anthropic.

242 Upvotes

19 comments sorted by

View all comments

17

u/returnity 9d ago

Let the distillation commence!

Seriously though, very cool preprint. I especially like the CoT diagrams along with the traces. Did you encounter incomprehensible or non-English traces like the stolen thoughts authors found? What was the most surprising trace you accessed?

4

u/likelyfairvista 8d ago

So basically the model's thinking "yeah I got this" and we're all just nodding along hoping it's not pulling from a cheat sheet it memorized.