Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Ataraxos Beat the World's Best Stratego Player 15-1 on 16 GPUs and About $8,000 — With a Fraction of DeepNash's Training Data

Ataraxos Beat the World's Best Stratego Player 15-1 on 16 GPUs and About $8,000 — With a Fraction of DeepNash's Training Data

A Nature paper from researchers at MIT, Carnegie Mellon, NYU and Stanford describes Ataraxos, an AI that beat the world's top human Stratego player 15-1 with four draws and surpassed DeepMind's DeepNash. Training used under 1/100 of DeepNash's samples and fewer than 1/30 of its self-play games.

A research team spanning MIT, Carnegie Mellon, NYU and Stanford has built an AI system that beat the world's best human Stratego player 15-1 — and it did so with a training budget that would not cover a week of a frontier model's electricity bill. The system, called Ataraxos, is described in a paper published in Nature.

Stratego is a genuinely hard testbed. Each player commands 40 pieces whose identities stay hidden until two pieces collide, and the number of possible board configurations exceeds 10^66 — vastly more than chess. A single game can stretch to roughly 2,000 moves, and the game rewards bluffing: move a scout as if it were a marshal, but not so often that your threats stop being believed.

Ataraxos was trained with self-play reinforcement learning, playing against itself repeatedly to build a base strategy, combined with more efficient training algorithms that avoid enumerating every possible opponent action. At decision time, it uses generative models to infer the probable identity of hidden enemy pieces from the current board state, then evaluates candidate moves against those beliefs. The team credits that decision-time planning as the key to surpassing human play.

The efficiency numbers are the story. Training used less than one-hundredth of the samples and fewer than one-thirtieth of the self-play games of DeepNash, DeepMind's 2022 system that had previously reached top-human level at Stratego. Ataraxos ran on 16 GPUs and cost roughly $8,000 — a rounding error next to the supercomputer budgets behind Deep Blue or AlphaGo.

Against DeepNash itself, Ataraxos came out ahead. Against the world's top human players it went 15-1 with four draws, and in world-championship play against elite humans it recorded 39 wins against 2 losses.

The researchers also report that applying the same approach to other imperfect-information games produced superhuman results, suggesting the recipe is general rather than Stratego-specific. The method borrows counterfactual regret minimization, a technique developed for poker AI, to balance risk and reward while bluffing just enough to keep an opponent guessing.

The point of the exercise is not the board game. Imperfect information is the normal condition of the real world: a trader does not know other participants' reasoning, a defender does not know an attacker's position, a negotiator does not know the other side's reservation price. A system that reaches superhuman decisions under hidden information, on an $8,000 budget, changes what is affordable for anyone building decision systems in those domains.

It also adds a data point to a shifting debate about AI progress. The frontier narrative has centered on scale — more parameters, more data, more compute. Ataraxos is the opposite argument: that algorithmic ideas, applied to a well-chosen problem, can still buy superhuman performance for the price of a used car.

Comments (0)

Log in to join the discussion

Log In

No comments yet