What it is
The authors introduce Ataraxos, an AI for the board wargame Stratego built on general techniques they developed for self-play reinforcement learning and for test-time search when much of the game is hidden. Ataraxos defeated the most decorated human Stratego player of all time by a large margin, to the authors' knowledge the first superhuman result in the game's history, while using orders of magnitude less compute and data than previous efforts. The same techniques produced a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all at low cost and with high sample efficiency.
Why it matters
Real-world decisions usually involve hidden information, known to one party but not to another, and large amounts of it make established reinforcement learning and search approaches ineffective. Top-human-level Stratego, a game with hidden information on a massive scale, had stayed beyond the reach of AI even with multimillion-dollar industrial research efforts. Because the approach worked across adversarial, cooperative and team games, the authors present it as a design pattern for reinforcement learning and search under large amounts of hidden information.
Every metric behind this entry is listed, with its source, under Sources and data below.
Filed underReinforcement Learning in Robotics, Artificial Intelligence in Games, Optimization and Search Problems