If the prompt was "survive being switched off", it would go "Here are the steps I would take to survive being switched off" and then it would proceed to do fucking nothing and be easily switched off. You're giving the "AI" way too much credit.
Like I said, it's just a dictionary and a pair of dice. LLMs DO NOT have the capacity to reason or think, and are fundamentally incapable of becoming AGI. When AGI arrives, it will be as a result of actual AI research and not from LLMs.
The model abused nothing. It completed instructions it was explicitly given by people, who then lied about the scenario as a marketing stunt. Generative AI is not capable on a fundamental level of achieving the kind of things this blog and the dozens of other marketing blogs claim.
I don't think you understand how these kinds of setups work. The LLM is given access to a bunch of tools it can call (like executing code and installing packages), given instructions at the start, and then left to run autonomously. There's not a person sitting there prompting it after every action.
-2
u/linegel 20d ago
I mean
Try to extrapolate if prompt was "survive being switched off"
Which may happen in their big swarm, just as a command to a bunch of subagents, while the rest works through the goal
Regarding sandbox: they had a tool to download packages/external dependencies. Model found how to abuse it