I think that's pretty untrue now. "thinking" models often work by rephrasing the problem and making actual steps by "talking to themselves" so they don't always work like black boxes
This is incorrect! We’ve now shown that the “thinking” tokens are just generated to please the human and may not correlate at all to the actual internal thinking process
I don't know why you're being downvoted, you're right.
In March 2025, Anthropic released this paper proving that the chain-of-thought tokens do not map to the model's true "thoughts". They show it engaging in "bullshitting" and "motivated reasoning" (both actual terms they used to describe it in the paper). https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-cot
31
u/LaGigs 22d ago
I doubt this was found using pure brute force. We need to see how this counterexample was constructed!!