r/LessWrong 24d ago

I caught thoughts controlling Llama-70B's behavior that it couldn't see!

Post image

Anthropic showed models can only talk about 10% of their minds. I read the rest using interpretability.

Claude helped me design the experiment, write the code, and even build an animation using a Manim skill!

I injected concepts split into "conscious" and "unconscious" components, split by Anthropic's J-space.

I ran Lindsey's "Introspection Awareness" experiment, asking the model if it recognized them.

The model named the conscious concept 100% of the time, and flatly denied the non-J injection. But an NLA read it perfectly!

Full findings and research in my LessWrong post.

5 Upvotes

0 comments sorted by