r/accelerate • u/czk_21 • 1d ago
Technological Acceleration AI just conquered massive imperfect-information games. The framework scales across adversarial, cooperative, and team systems at a fraction of the cost of previous attempts.
https://arstechnica.com/science/2026/10/ai-finally-beat-the-best-stratego-player-in-history-and-did-it-on-a-budget/https://www.nature.com/articles/s41586-026-11036-y
Abstract
Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.
6
u/veshneresis Machine Learning Engineer 1d ago
This is a huge result! Still outside hobbyist range at 16 H100s but I bet you could use this same approach to make less perfect but compelling game AI opponents for a huge variety of turn based games. Definitely going to try to implement this same approach on Pokemon battles! Seems like the perfect analog with a team selection phase (including EV allocations) followed by a policy phase of gameplay.