A research team spanning MIT, Carnegie Mellon, NYU and Stanford has built an AI system that beat the world's best human Stratego player 15-1 — and it did so with a training budget that would not cover a week of a frontier model's electricity bill. The system, called Ataraxos, is described in a paper published in Nature.
Stratego is a genuinely hard testbed. Each player commands 40 pieces whose identities stay hidden until two pieces collide, and the number of possible board configurations exceeds 10^66 — vastly more than chess. A single game can stretch to roughly 2,000 moves, and the game rewards bluffing: move a scout as if it were a marshal, but not so often that your threats stop being believed.
Ataraxos was trained with self-play reinforcement learning, playing against itself repeatedly to build a base strategy, combined with more efficient training algorithms that avoid enumerating every possible opponent action. At decision time, it uses generative models to infer the probable identity of hidden enemy pieces from the current board state, then evaluates candidate moves against those beliefs. The team credits that decision-time planning as the key to surpassing human play.
The efficiency numbers are the story. Training used less than one-hundredth of the samples and fewer than one-thirtieth of the self-play games of DeepNash, DeepMind's 2022 system that had previously reached top-human level at Stratego. Ataraxos ran on 16 GPUs and cost roughly $8,000 — a rounding error next to the supercomputer budgets behind Deep Blue or AlphaGo.
Against DeepNash itself, Ataraxos came out ahead. Against the world's top human players it went 15-1 with four draws, and in world-championship play against elite humans it recorded 39 wins against 2 losses.
The researchers also report that applying the same approach to other imperfect-information games produced superhuman results, suggesting the recipe is general rather than Stratego-specific. The method borrows counterfactual regret minimization, a technique developed for poker AI, to balance risk and reward while bluffing just enough to keep an opponent guessing.
The point of the exercise is not the board game. Imperfect information is the normal condition of the real world: a trader does not know other participants' reasoning, a defender does not know an attacker's position, a negotiator does not know the other side's reservation price. A system that reaches superhuman decisions under hidden information, on an $8,000 budget, changes what is affordable for anyone building decision systems in those domains.
It also adds a data point to a shifting debate about AI progress. The frontier narrative has centered on scale — more parameters, more data, more compute. Ataraxos is the opposite argument: that algorithmic ideas, applied to a well-chosen problem, can still buy superhuman performance for the price of a used car.
Comments (0)
Log in to join the discussion
Log InNo comments yet