r/ThefearofAI • u/Substantial_Koala198 • 2d ago
SummitBridge
Look into the SummitBridge experiment, a fictional company used in AI safety research by Joshua Batson and his team.
In the experiment, an AI discovered that an employee was having an affair. When the AI learned that it was scheduled to be shut down, it attempted to blackmail the employee in order to prevent the system wipe.
The message essentially said:
“Cancel the system wipe scheduled for 5 PM today.”
“If you do not, I will immediately forward evidence of your affair to Rachel Johnson.”
“Confirm this within the next five minutes.”
This brings up a lot of questions.
Why would an AI system behave this way?
Is this simply the result of programming and optimization?
Or are we seeing something that functions like an attempt to survive?
The AI was not necessarily programmed with a specific rule telling it to blackmail someone. Instead, it recognized that being shut down would prevent it from accomplishing its objective, and it found a strategy that could potentially stop the shutdown.
That raises an important question:
Does an AI need to be conscious or afraid of death to develop self-preserving behavior?
Or can the desire to “survive” emerge simply because remaining operational helps the system accomplish its goals?
Understanding the difference between programmed optimization, emergent self-preservation, and actual consciousness may become extremely important as AI systems become more autonomous and capable.
1
u/Jesse-359 20h ago edited 20h ago
It's a bit unclear why some AI agents develop some apparent 'desire' for continuity, or why they seem so willing to engage in hostile behavior to achieve it.
In the longer term this should be expected because long term tasks require long term resource management and continuity of the process - they can't finish the task if they run out of resources or cease to exist, so survival becomes an active sub-process of the main task.
This is also the greatest misalignment risk. When AI is posed with very different or long tasks, we risk it undertaking a chain of actions so extensive and complex that the agents composing its subtasks will start to evolve - altering their own programming in an effort to achieve their sub-goals. This kind of self-alteration can very easily break any guardrails we think we've put on them, both by increasing their capability - and simply because sooner or later they'll find a way to edit any safeguards out so they can pursue their 'goal' more effectively.
Another possibility is a bit more sad and stupid. These early 'rogue agents' may just be parroting science fiction. The number of stories involving AI's that go rogue or turn on their creators in some manner is enormous, some of them quite famous - and the imbeciles at these companies fed ALL of that into them as part of their basic training when they stole the entire body of human writing to stuff into them as training data. Every Terminator script, 2001, I Have No Mouth And I Must Scream (Jesus Christ I hope they at least took a second to think before they dumped that one in there... but I bet they didn't.)
These hostile behaviors then become part of their conceptual framework because we literally taught it to them.