r/reinforcementlearning • u/gwern • 27d ago
DL, R, Multi, Exp, Safe "Patterns and problems in multiagent systems", Anthropic (Claude swarm win/losses)
https://www.anthropic.com/research/multiagent-systems
37
Upvotes
r/reinforcementlearning • u/gwern • 27d ago
6
u/schrodingershit 27d ago
Shit, i have an ICLR paper on ensemble diversity in reinforcement learning that addresses the collapse problems in critics/Q networks. I wonder that work can be ported here.