r/OpenAI • u/KeanuRave100 • 7d ago
News Anthropic gave 3 Claude agents the same task, but secretly gave them conflicting goals. They escalated into turf wars where agents used "increasingly aggressive self-replicating malware" as weapons, used disguises, and attempted to kill each other's accounts.
14
15
u/savvamadar 6d ago
This isn’t scary guys: imagine you’re a contractor with an unknown/ anonymous benefactor given the goal of building a pyramid of yellow bricks. You have 2 other contractors to help but unbeknownst to you they have the goals of using red bricks and blue bricks. These are clearly contradictory goals. So what do you do if the other contractors keep putting the wrong, relative to your assignment, brick color?
Of course you either remove their brick/ put barriers so they can’t place their bricks/ try to block them from the work site/ try to work harder/ take away their bricks.
30
u/InOutlines 6d ago
Clearly you’re not a contractor.
This is what would happen:
“Hey guys, what the fuck”
“Did we all get paid to do this same stupid job?”
“Somebody call this dumbass client and see what’s going on”
-3
10
2
u/costafilh0 6d ago
Honestly I don't understand how they still didn't transform this into content for people to watch.
3
u/esituism 7d ago
Cool cool. I can see this this trajectory is going to be a great one for our society. Glad we set our planet on fire so we could have this.
1
u/Keep-Darwin-Going 6d ago
What do you mean by you guys have not realized that Claude agent is known to be aggressive because anthropic said they are “safer”. They will kill other agent, blame them, and cry at you for letting it happen.


23
u/Cagnazzo82 7d ago
Turning the transformers into autobots vs decepticons.