Most agent benchmarks grade what an AI assistant says. A new one from Microsoft grades what it leaves behind in your database. ThinkingBox, published on the Hugging Face blog on October 3 by Microsoft's Tuhin Kundu with co-authors from Hugging Face, Enderis AI and researchers at the University of Pittsburgh, UC Irvine and Northwestern, runs agents through 507 synthetic enterprise workflows — retail, auto insurance, travel, a neobank, consulting — and then inspects the terminal backend state: the records changed, the side effects, the fields that should have changed and didn't.
The numbers are sobering. In an ablation covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks — and 67.24% of those failures terminated cleanly, invoked a state-changing tool, and reported no error at all. The authors' framing is three sentences long: "A trajectory is a claim. Database state is the evidence. Repetition is the trust test."
Among the failed runs, executable checks found wrong field values in 77.61% of cases, unintended extra effects in 43.30%, and missing required changes in 25.36% — categories that overlap. Roughly four in five failures were classified as tool-handling problems rather than reasoning flaws: agents failing to recover from empty lookups or unmet preconditions, not misunderstanding the task.
The benchmark's canonical example is a $745 kitchen appliance sitting fifteen days late in a Nashville distribution center. The agent makes nine valid tool calls, pulls the order, reads the refund policy twice, opens a ticket — then sets its status to "solved" and tells the customer everything is resolved. The carrier exception is still open; procedure required "hold." The customer never got an answer. Grading the transcript catches nothing. Only reading the status field shows the failure.
ThinkingBox also separates capability from dependability. Each task runs 20 times from an identical clean backend, producing three metrics: pass@1, pass@20 (solved at least once), and observed 20/20 (solved every single time). On pass@1, Claude Opus 5.5 leads at 67.16%, ahead of Claude Opus 5 at 66.50% and GPT-5.4 at 65.36%, with GPT-6 Astra at 58.31%. The strongest open-weight model, Kimi-K3, hits 57.37% — but solves 93.89% of tasks at least once and only 68 of 507 (13.41%) in all twenty runs. Claude Opus 5 and Opus 5.5 each passed exactly the same 241 tasks on every attempt, meaning Opus 5.5's higher headline score bought zero additional dependable tasks. GPT-6 Astra kept 78% of its single-attempt score across the 20-run series; GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro kept roughly 8%.
Priced at undiscounted OpenRouter list rates from a September 20 snapshot, dependability has a bill attached: GPT-5.4 was cheapest per dependable task at $6.80 (128 tasks passing all 20 runs), GPT-6 Astra at $7.45 (231 tasks), Claude Opus 5.5 at $7.80 (241), and GPT-5.6 Sol — the cheapest per single success at $0.127 — rose to $9.76 once reliability was priced in. Domain variance is extreme: Claude Opus 4.6 scored 68.62% on retail workflows and 8.30% on auto insurance.
The harness and benchmark, ThinkingBox-Bench, are open-sourced through OpenEnv — code under MIT, data under CDLA-Permissive-2.0. The authors recommend checking terminal state before committing changes, classifying tool errors so retries target recoverable ones, trimming the exposed tool surface, and requiring human approval for changes that cannot be cheaply reversed. They are careful to note they have not yet measured the lift from any of those steps. An RL training repository is listed as coming soon. For teams shipping agents into production this fall, the message is blunt: your monitoring watches for errors, and the failures that matter are the ones that never raise one.
Comments (0)
Log in to join the discussion
Log InNo comments yet