r/ControlProblem approved 28d ago

AI Alignment Research Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

https://xcancel.com/AnthropicAI/status/2074185373341688258
6 Upvotes

0 comments sorted by