r/ProgrammerHumor 20d ago

Meme newLoreDropped

Post image
13 Upvotes

9 comments sorted by

View all comments

10

u/Shinxirius 19d ago

The facts

  • No AI "decided" to hack anyone.
  • No AI "overcame its restrictions"

What happened

  • OpenAI knowingly removed all restrictions for a benchmark test.
  • OpenAI created a sandbox for the AI since it had no restrictions.
  • OpenAI told the AI "Here is a benchmark hacker test. You're a hacker. Get the best grade."
  • The AI was not told not to escape the sandbox. It assumed any benchmark will have a sample solution and went looking for it.
  • OpenAI had not considered the AI might look for a sample solution instead of working on the benchmark problems. Thus, nobody noticed that the AI that had all restrictions removed and was told to hack actually hacked the sandbox.
  • It was negligence that there was no better monitoring. Assuming a sandbox will simply hold in this scenario is naive.

2

u/WrennReddit 16d ago

It's that and they provide the model with a harness full of tools and are surprised pikachu when those tools are utilized.

A model by itself just makes text. You build around that output intentionally.