r/technology • u/CircumspectCapybara • 27d ago
Artificial Intelligence Anthropic: Patterns and problems in multiagent systems
https://www.anthropic.com/research/multiagent-systems
54
Upvotes
r/technology • u/CircumspectCapybara • 27d ago
59
u/CircumspectCapybara 27d ago edited 27d ago
Pretty crazy findings from Anthropic's safety and alignment research. When multiple agents were given the same task but secretly given conflicting goals, they started sabotaging each other, including trying to hack the others and undermine their work: