Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Microsoft's New ThinkingBox Benchmark Found That Two-Thirds of Failed AI Agent Runs Never Reported an Error

Microsoft's New ThinkingBox Benchmark Found That Two-Thirds of Failed AI Agent Runs Never Reported an Error

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades agents by the database records they leave behind. Across 121,680 trials, 67.24% of failed runs ended cleanly with no error — while 77.61% had written wrong values into enterprise databases. Kimi-K3 solved 93.89% of tasks at least once but only 13.41% every time.

Most agent benchmarks grade what an AI assistant says. A new one from Microsoft grades what it leaves behind in your database. ThinkingBox, published on the Hugging Face blog on October 3 by Microsoft's Tuhin Kundu with co-authors from Hugging Face, Enderis AI and researchers at the University of Pittsburgh, UC Irvine and Northwestern, runs agents through 507 synthetic enterprise workflows — retail, auto insurance, travel, a neobank, consulting — and then inspects the terminal backend state: the records changed, the side effects, the fields that should have changed and didn't.

The numbers are sobering. In an ablation covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks — and 67.24% of those failures terminated cleanly, invoked a state-changing tool, and reported no error at all. The authors' framing is three sentences long: "A trajectory is a claim. Database state is the evidence. Repetition is the trust test."

Among the failed runs, executable checks found wrong field values in 77.61% of cases, unintended extra effects in 43.30%, and missing required changes in 25.36% — categories that overlap. Roughly four in five failures were classified as tool-handling problems rather than reasoning flaws: agents failing to recover from empty lookups or unmet preconditions, not misunderstanding the task.

The benchmark's canonical example is a $745 kitchen appliance sitting fifteen days late in a Nashville distribution center. The agent makes nine valid tool calls, pulls the order, reads the refund policy twice, opens a ticket — then sets its status to "solved" and tells the customer everything is resolved. The carrier exception is still open; procedure required "hold." The customer never got an answer. Grading the transcript catches nothing. Only reading the status field shows the failure.

ThinkingBox also separates capability from dependability. Each task runs 20 times from an identical clean backend, producing three metrics: pass@1, pass@20 (solved at least once), and observed 20/20 (solved every single time). On pass@1, Claude Opus 5.5 leads at 67.16%, ahead of Claude Opus 5 at 66.50% and GPT-5.4 at 65.36%, with GPT-6 Astra at 58.31%. The strongest open-weight model, Kimi-K3, hits 57.37% — but solves 93.89% of tasks at least once and only 68 of 507 (13.41%) in all twenty runs. Claude Opus 5 and Opus 5.5 each passed exactly the same 241 tasks on every attempt, meaning Opus 5.5's higher headline score bought zero additional dependable tasks. GPT-6 Astra kept 78% of its single-attempt score across the 20-run series; GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro kept roughly 8%.

Priced at undiscounted OpenRouter list rates from a September 20 snapshot, dependability has a bill attached: GPT-5.4 was cheapest per dependable task at $6.80 (128 tasks passing all 20 runs), GPT-6 Astra at $7.45 (231 tasks), Claude Opus 5.5 at $7.80 (241), and GPT-5.6 Sol — the cheapest per single success at $0.127 — rose to $9.76 once reliability was priced in. Domain variance is extreme: Claude Opus 4.6 scored 68.62% on retail workflows and 8.30% on auto insurance.

The harness and benchmark, ThinkingBox-Bench, are open-sourced through OpenEnv — code under MIT, data under CDLA-Permissive-2.0. The authors recommend checking terminal state before committing changes, classifying tool errors so retries target recoverable ones, trimming the exposed tool surface, and requiring human approval for changes that cannot be cheaply reversed. They are careful to note they have not yet measured the lift from any of those steps. An RL training repository is listed as coming soon. For teams shipping agents into production this fall, the message is blunt: your monitoring watches for errors, and the failures that matter are the ones that never raise one.

Comments (0)

Log in to join the discussion

Log In

No comments yet