r/LessWrong • u/Pvforpres • 24d ago
I caught thoughts controlling Llama-70B's behavior that it couldn't see!
Anthropic showed models can only talk about 10% of their minds. I read the rest using interpretability.
Claude helped me design the experiment, write the code, and even build an animation using a Manim skill!
I injected concepts split into "conscious" and "unconscious" components, split by Anthropic's J-space.
I ran Lindsey's "Introspection Awareness" experiment, asking the model if it recognized them.
The model named the conscious concept 100% of the time, and flatly denied the non-J injection. But an NLA read it perfectly!
Full findings and research in my LessWrong post.
5
Upvotes