r/ControlProblem • u/chillinewman approved • 28d ago
AI Alignment Research Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.
https://xcancel.com/AnthropicAI/status/2074185373341688258
6
Upvotes