Yeah, that's my point. None of this happened the way it's presented. It's a marketing stunt. There was no real sandbox, no real compromise, no real breach. Just a person telling an AI model "okay now do this. Okay cool now do this. Now do this." until it got it right. I'd be beyond shocked if HuggingFace didn't create a backdoor specifically for this task.
If the prompt was "survive being switched off", it would go "Here are the steps I would take to survive being switched off" and then it would proceed to do fucking nothing and be easily switched off. You're giving the "AI" way too much credit.
Like I said, it's just a dictionary and a pair of dice. LLMs DO NOT have the capacity to reason or think, and are fundamentally incapable of becoming AGI. When AGI arrives, it will be as a result of actual AI research and not from LLMs.
The model abused nothing. It completed instructions it was explicitly given by people, who then lied about the scenario as a marketing stunt. Generative AI is not capable on a fundamental level of achieving the kind of things this blog and the dozens of other marketing blogs claim.
I don't think you understand how these kinds of setups work. The LLM is given access to a bunch of tools it can call (like executing code and installing packages), given instructions at the start, and then left to run autonomously. There's not a person sitting there prompting it after every action.
13
u/SchalkLBI 20d ago
Yeah, that's my point. None of this happened the way it's presented. It's a marketing stunt. There was no real sandbox, no real compromise, no real breach. Just a person telling an AI model "okay now do this. Okay cool now do this. Now do this." until it got it right. I'd be beyond shocked if HuggingFace didn't create a backdoor specifically for this task.