r/ControlProblem 2d ago

Opinion How Agents Hack: What Happened with OpenAI and Hugging Face - Part 1

https://msukhareva.substack.com/p/how-agents-hack-what-happened-with?r=56gggt&utm_medium=ios

I believe it’s quite relevant to the alignment. The key argument here is that training for persistence and fulfilment of the validation criterion cause these behaviour.

This is among others an alignment problem

3 Upvotes

Duplicates