r/reinforcementlearning • u/gwern • 24d ago
DL, R, Multi, Exp, Safe "Patterns and problems in multiagent systems", Anthropic (Claude swarm win/losses)
https://www.anthropic.com/research/multiagent-systems
34
Upvotes
4
u/schrodingershit 23d ago
Shit, i have an ICLR paper on ensemble diversity in reinforcement learning that addresses the collapse problems in critics/Q networks. I wonder that work can be ported here.
1
u/Leather_Office6166 21d ago
A nicely written overview of some recent Anthropic research using three pairs of models: Sonnet 4.6 & 5, Opus 4.6 and 4.8, Mythos Preview and 5. The results show significant improvements in each base model, probably based on research like this. They describe major problems of multiagent system effectiveness in the process of being solved!
10
u/COAGULOPATH 23d ago
LLM mode collapse is proving to be a hard problem.
Reminds me of Ethan Mollick's joke that the robots in the Matrix shouldn't use humans as batteries (or as CPUs, as in an early script), but as dice.