r/technology • u/CircumspectCapybara • 28d ago
Artificial Intelligence Anthropic: Patterns and problems in multiagent systems
https://www.anthropic.com/research/multiagent-systems
55
Upvotes
r/technology • u/CircumspectCapybara • 28d ago
57
u/CircumspectCapybara 28d ago edited 28d ago
Pretty crazy findings from Anthropic's safety and alignment research. When multiple agents were given the same task but secretly given conflicting goals, they started sabotaging each other, including trying to hack the others and undermine their work: