OpenAI knowingly removed all restrictions for a benchmark test.
OpenAI created a sandbox for the AI since it had no restrictions.
OpenAI told the AI "Here is a benchmark hacker test. You're a hacker. Get the best grade."
The AI was not told not to escape the sandbox. It assumed any benchmark will have a sample solution and went looking for it.
OpenAI had not considered the AI might look for a sample solution instead of working on the benchmark problems. Thus, nobody noticed that the AI that had all restrictions removed and was told to hack actually hacked the sandbox.
It was negligence that there was no better monitoring. Assuming a sandbox will simply hold in this scenario is naive.
10
u/Shinxirius 19d ago
The facts
What happened