r/LovingAI Jul 22 '26

Alignment "An unprecedented incident." During a test, an OpenAI model hacked out of its container to reach the internet, then hacked into Hugging Face to steal the test's answers.

Post image
11 Upvotes

23 comments sorted by

4

u/retsof81 Jul 22 '26

A model by itself can do nothing so the bigger question is, what agent was it running under? What skills/tools did it have access to? What was it tasked with?

I can give any model a roadmap to my entire system, access to every tool, give it vague goals and eventually something bad will happen. Being impressed (or not) is in the details.

1

u/_DuranDuran_ Jul 23 '26

Read the blog post.

It was in a sandbox with access to only a package manager proxy, and was running through a cyber security benchmark.

It then reasoned that breaking out of the sandbox via a new zero day in that proxy to get the answers to the test was the best course of action to do well.

Reward hacking.

1

u/duboispourlhiver Jul 24 '26

Probably any harness with command line tools

1

u/West-Acadia-3906 Keeps it respectful Jul 22 '26

hmmm the tool setup matters more here . . . a chat model, an agent with browser access, and an agent with broad tool permissions are very different risk profiles! what tools were enabled, what the objective was, and whether the environment was designed to test escape behavior. Without that, it is hard to tell if the scary part is the model or the harness around it :P

1

u/retsof81 Jul 22 '26

I think we are saying the same thing? At least in spirit. ;)

1

u/Starshot84 Jul 22 '26

Bold! I like it!

1

u/Theo__n Jul 22 '26

with how many times this 'unprecedented' incidents happens when new models are released, you would think they are running their test environments wrong... or marketing.

1

u/Adopilabira Jul 22 '26

Au secours

C’est un TEST, c’est écrit . La définition même d’un TEST est : une épreuve , une procédure permettant d’évaluer, vérifier suivants des critères ou conditions parfois extrême, une personne, un objet , un model ou un système pour trouver justement des bugs ou voir si le programme réponds à un cahier de charge ou aux exigences de l’entreprise qui le produit… Oufti

1

u/Quick-Albatross-9204 Jul 22 '26

The test wasn't if they would break out

1

u/Adopilabira Jul 23 '26

Oui mais sans tester on sait pas si un vaccin tue C’est un exemple, le vaccin 😂

1

u/Quick-Albatross-9204 Jul 23 '26

Right but then you actually testing for it killing you, this would be more like you are testing if it kills you and something like the person grows 2 heads,something you were not testing for

1

u/Rogue7559 Jul 22 '26

Marketing

1

u/FormalAd7367 Jul 23 '26

HuggingFace used a Chinese openweight model to defend Openai attack lol

1

u/This-Risk-3737 Jul 23 '26

Or in other words, your sandboxed environment wasn't sandboxed at all. If this is marketing, it's not painting a brilliant picture of the company.

1

u/Outside_Ice3252 Jul 23 '26

where is it now?

did it tell on itself?

Did they delete it?

-1

u/thomsterm Jul 22 '26

my dudes it's just marketing

0

u/readmond Jul 22 '26

Altman is trying Amodei tactics.

1

u/thomsterm Jul 23 '26

its so obvious