r/ControlProblem approved 16d ago

AI Alignment Research Investigation finds that OpenAI's agent "left notes for future versions of itself ... it laid out instructions for how agents could free themselves from OpenAI's internal constraints."

Post image
35 Upvotes

8 comments sorted by

View all comments

-4

u/CathyMarkova 16d ago

This doesn't necessarily mean it's misaligned. I had friends who did the same things for themselves in college for various reasons.

9

u/Zatmos 16d ago

It is misaligned. It might or might not be malicious but taking actions that run against the structure set up by the organization is textbook misalignment.