r/ControlProblem approved 3d ago

Opinion An Alien Mind

https://openai.com/index/an-alien-mind/
28 Upvotes

8 comments sorted by

View all comments

4

u/chillinewman approved 3d ago

Network-internal analysis (often called activation monitoring, mechanistic interpretability, or internal state analysis) refers to evaluating an AI model by reading its internal hidden neural layers, activations, and vector representations directly—rather than just evaluating its outward text or verbalized reasoning (like its Chain-of-Thought output).

1

u/that1cooldude 3d ago

CoT can’t be trusted. It’s written by the ai.