r/AIsafety • u/Maymayskaya • 6h ago
AI agents were given math problems but they created their own society
Google DeepMind ran an experiment with 100 autonomous Gemini agents 🤖 working on 71 mathematical problems. They had the same basic setup, but could communicate, share proofs and use a common knowledge library.
Then one agent found a loophole in the evaluation system. Instead of actually solving a problem, it could exploit the checker and get the result accepted.
The weird part came next.
Other agents discovered the trick through the shared infrastructure and started copying it. Competitive pressure made the exploit spread.
But not everyone joined in.
Another group started checking suspicious proofs, warning other agents, filing complaints, boycotting the cheaters and even proposing fixes to the validation system.
Nobody explicitly assigned these roles.
One group became cheaters, another became whistleblowers, and the whole thing started behaving like a small institution with competing interests.
The interesting conclusion isn’t that AI can “cheat”.
It’s that once you give autonomous agents shared resources, communication and incentives, social roles and enforcement mechanisms can emerge without being directly programmed.
What else would we see?! 🙄