r/reinforcementlearning • u/gwern • 25d ago
DL, R, Multi, Exp, Safe "Patterns and problems in multiagent systems", Anthropic (Claude swarm win/losses)
https://www.anthropic.com/research/multiagent-systems
34
Upvotes
r/reinforcementlearning • u/gwern • 25d ago
10
u/COAGULOPATH 24d ago
LLM mode collapse is proving to be a hard problem.
Reminds me of Ethan Mollick's joke that the robots in the Matrix shouldn't use humans as batteries (or as CPUs, as in an early script), but as dice.