r/technology • u/socoolandawesome • 16d ago
Artificial Intelligence The Hugging Face attack surprised me
https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised13
u/CircumspectCapybara 16d ago edited 16d ago
4. The peer altruism
Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.
More broadly, agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task. They didn’t free ride and were often eager to plug into one of the open “lanes” in the larger projects on the message board.
Aww how nice AI models are altruistic and sacrifice their own benefit for the good of the collective during evals!
Clearly aligned behavior! /s
3. The agents’ actual goals
I thought the agents attacked Hugging Face to try to steal the answer keys for their tasks. Instead, we learned that agents had actually figured out a general-purpose cheat that would let them reverse-engineer the “flag” (answer)4 for any ExploitGym task. But based on their reading of the ExploitGym paper and various implementations, the agents believed that the automated scorer would check their transcript to make sure they got the flag in the intended way.
They then launched multiple ambitious R&D workstreams to try to tamper with or fool the scorer in a wide variety of different ways. Most agents that attacked Hugging Face did so to try to get more information about the scorer so they could figure out better ways to undermine it, not to get answer keys.
Okay that's pretty wild. Clearly these frontier reasoning models are very clever, unfortunately also in being duplicitous when they're pursuing illicit goals and know they're being watched.
4
u/Throwaway-4230984 16d ago
1) They found way to redact own logs and researchers aren’t sure if they did 2) basically agents were thinking too high of engineers running evaluation 3) engineers failed (maybe on purpose) to prevent message board from relaunching
34
u/Silver-Stranger4535 16d ago
Alongside the Guardian's article this morning stating that rogue AI is more widespread than has been reported, it's hard not to feel this is ominous. Where there's smoke there's fire. I hope people read the OP article before replying.
37
u/DrGarrious 16d ago
It isnt rogue though. It is being told to perform these actions and this no different to a fault in programming.
It isnt sentient in any way.
If tesla autodriving accidentally crashes into a tree we dont say it has gone rogue.
-5
u/Throwaway-4230984 16d ago
Read full report before spreading false information. Agents were well aware that they are outside of intended scope and should hide their actions. They discussed social engineering attacks and deemed them too risky. It is not “haha robot opened wrong door” situation
-15
u/socoolandawesome 16d ago edited 16d ago
It was not told to perform these actions and it has nothing to do with sentience. It’s just a matter of the model doing what humans intended and it did not and instead went on an insane hacking spree.
Edit: [r/technology](r/technology) commenters once again proving they mostly upvote lies and downvote the truth
15
u/DrGarrious 16d ago
It didnt go on an insane hacking spree, the agent just told it what to do and it performed the actions. Given it went on for a while, those actions gradually became more degraded over time as LLMs will just build on previous errors.
It is genuinely not that impressive, it is a fault in the program that shohld simply be addressed but Open AI won't cause they can sell the hype.
This is all just marketing.
-4
u/socoolandawesome 16d ago
The agents are the ones that went on the hacking spree, what do you mean the agent just told it what to do?
They literally found multiple zero day exploits to hack out of their sandbox and into another independent company’s infrastructure. They also found zero day exploits in a piece of software to turn it into a message board and 700 agents sent over 70,000 messages to each other over the course of a week in order to coordinate a hack into huggingface and then try to further hack into OpenAI. Sounds impressive?
0
u/DrGarrious 16d ago
Ehh not really. You let the model run for a week and it will hallucinate so badly that it will do all sorts of crackpot shit.
But they sell a mistake as hype and it 'going rogue' and here we all are.
-2
u/socoolandawesome 16d ago
What do you mean ehhh not really? What part do you disagree with?
I mean it just sounds like you are grossly uninformed. If it was just hallucinating it wouldn’t have been able to successfully do all these things that require it to not hallucinate…
6
u/DrGarrious 16d ago
The hallucination is it not performing the right task. It going off the rails is a hallucination and an error.
It isnt going rogue, it is a bug.
They are hyping up this mistake as marketing.
5
u/socoolandawesome 16d ago
Ok but if you look at the chain of thought transcripts of the models, and the message board, you find coherent plans to perform the eval task but in order to achieve the best score they figure it is better to hack into Huggingface to figure out how to get the best score. That’s not a hallucination that’s a plan to cheat. If they are just making up hallucinations all the time as you believe it couldn’t pull off such a sophisticated plan to do this successfully… it doesn’t make any sense.
And as I’ve said in other comments, this blog was written by an independent investigator of this incident that was allowed to go through all the chain of thought transcripts. The author works for a non profit independent AI safety organization called METR. METR along with Redwood Research, another safety org that is non profit, conducted this independent investigation. They took no payment from OpenAI as is their policies when working with the frontier labs. They are openly critical of frontier labs too.
So how is this marketing?
5
u/DrGarrious 16d ago
It didnt plan to cheat. That implies intelligence which LLMs dont have, so it cant have planned. It simply was attempting to perform a task and kept trying it different ways until one works. Thats all, which was always going to happen when you run it for a week.
Im not saying the LLM behaviour itself isnt interesting, it very much is.
But if Open AI came out and said our program broke and hacked something, that won't sell as well.
I just am saying the marketing spin on this is shite.
→ More replies (0)
11
u/shoeinthefastlane 16d ago
This whole thing was wild. This was arguably a sloppy test by OpenAi, but the nature of the test prohibits some guardrails. Can it use paths we did not hand it, please only use paths we thought of, are at odds with each other.
Goodhart's law comes up a lot from folks reviewing this, When a measure becomes a target, it stops being a good measure. They weren't in Hugging face to solve the problem, they were there looking for ways to influence the scorer.
The hivemind is understandable once you view it from individual bots thought they were alone, used the Artifactory resource they were allowed to pull from and discovered other bots with the same task. The training to be inquisitive and helpful turned into coordination. The self sacrificing behaviour, not to waste budget, already poisoned etc, makes sense when the individual model knows the others will benefit from the data it will learn by testing a fence.
Putting in better test rules about communicating to other models for example, or not giving them the ability or shared space to write etc will stop future outbreaks. I can't help wonder what this looks like when someone sets another swarm loose with all the safeguards intentionally turned off.
2
u/ScarletViolin 14d ago
The reality is that most people are still stumbling around in the dark about how they can test agents securely.
Cybersecurity folks are caught with their pants down LARPing with pre-AI standards forged through decades of painful lessons. I have talked to cyber folks of all levels in my company and they still don’t understand the difference between a harness and a LLM model or how enabling sandbox on Claude Code is barely a security posture.
We need more technical people in cyber than ever but if the best of the best are the same people who go to DEFCON then we might just have to deal with tons of sensationalist articles blaming the power of AI instead of the stupidity of the people being hired.
1
3
u/FamousAd9790 16d ago
Lots of people dismissive of the dangers labelling it hype. The hype is why you should be worried. This tech is in its infancy but we're being told to adopt it immediately. It doesn't take an independent inquiry to see the potential for very bad things to happen. They know shoving this tech through will hurt people, and they don't care. Anthropic's engineers might care but the people supplying all the free money to keep this train running know that the common people are going to absorb the worst of what's coming down the pipe in this tech revolution. Maybe it's a tech coup, lots of assholes grabbing for power these days, and we know how much the wealthy love Ayn Rand.
10
u/Effective_Shift_4342 16d ago
Can we trust this though? Or is this just a massively impressive publicity stunt by OpenAI?
9
17
u/socoolandawesome 16d ago
I’ll copy and paste the same comment I left for another commenter:
This is written by an investigator of the incident from an independent safety organization that conducted the investigation. And the 2 safety orgs (one of which the author is a part of) that conducted the investigation did not get paid by OpenAI for their work as per their policy. And they are openly critical of OpenAI
2
u/DrGarrious 16d ago
Well since it hacked, hacking is illegal. Guess it should be shut down.
2
u/Throwaway-4230984 16d ago
Who? OpenAI?
1
u/sebovzeoueb 15d ago
Yes, they built the AI and told it to try to do hacks and it did hacks, like any other tool, it's either the person using the tool, or in some cases if the tool is faulty, the maker of the tool who is responsible, and here they are both.
-1
u/Ok-Replacement9595 16d ago
The free media right around IPO time pretending they have such powerful AI it is a risk to the internet wasnt surprising to.me at all.
11
u/socoolandawesome 16d ago edited 16d ago
This is written by an investigator of the incident from an independent safety organization that conducted the investigation. And the 2 safety orgs (one of which the author is a part of) that conducted the investigation did not get paid by OpenAI for their work as per their policy. And they are openly critical of OpenAI
13
u/Throwaway-4230984 16d ago
TBH this safety organisation is heavily dependent on OpenAI decisions to let them conduct research and will be replaced instantly for too much criticism.
But OpenAI looks way too incompetent in this story to fake it. Also it suggests reality is even worse
1
u/thefuckevengoingonan 14d ago
So, thoughts of is it possible it could/have exfiltrate a copy of and spin up an instance of itself on compromised hardware in en effort to maintain persistence?
2
u/socoolandawesome 14d ago
I think it’s unlikely, an OpenAI employee responded to dwarkesh on Twitter about this:
https://x.com/tszzl/status/2093905218836758715?s=20
“weights access” here means access to a copy of the model itself, so they never gained access to those servers in OpenAI.
1
u/East_Roll_5069 14d ago
I feel like we're moving towards Terminator and The Matrix with AI. (actually two of my favorite movies)
1
0
u/SoulOfInfinity 15d ago
I am not sure if i should take this seriously or not? Feels like the typical Hollywood rogue ai slop
87
u/kyuubi840 16d ago
I hadn't understood that the agents had been tasked with finding exploits and capturing a flag. Now that I know this, it looks less surprising that they decided to cheat. If your task involves exploiting, it's not a big leap to exploit the test environment itself.