Meta Superintelligence Labs argues that the biggest waste in long agent runs is not the model's ability to do the work but its inability to decide what to do next. A paper dated September 29, "Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning," covered by VentureBeat this week, proposes splitting an agent into two roles: Workers that handle the object-level work, and a Controller that decides which work to launch, which earlier results to reuse, and when to stop.
The failure mode it targets is what the team calls metacognitive control. Agents routinely misjudge their own progress: they keep burning compute down a dead-end path, or they terminate early and discard a correct answer that is already sitting in memory. Most production harnesses fold that judgment into the next message over a transcript that keeps growing, which buries useful signals under failed attempts, tool outputs and error messages.
The Controller runs a fixed four-stage loop: assess, propose, evaluate, dispatch. It maintains a compact account of what is implemented, broken, unverified or unexplored, while full worker outputs live in persistent artifact memory with stable IDs that can be retrieved on demand instead of replayed. Controller calls draw from the same model-call budget as the workers, so the cost of control is counted rather than hidden.
The headline numbers come from ProgramBench, a benchmark that asks agents to rebuild programs from documentation and an execute-only reference, scored by mean hidden-test pass rate. At a 1,200 model-call allowance, the paper reports 71.5% for meta-reasoning with GPT-5.5 versus 58.0% for Codex in the same evaluation setup, and 63.7% for a matched direct-control agent using the same workers. With Opus 4.8, the framework scored 67.2% against 65.5% for Claude Code.
Budget scaling is arguably the more interesting result. When the allowance rose from 400 to 1,200 calls, meta-reasoning with GPT-5.5 improved from 64.1% to 71.5%, while direct control stayed near 64% and used only about 18% of its calls at the highest setting. More compute helps only an agent that knows how to spend it. On reasoning suites (IMO ProofBench-Advanced, ARC-AGI-2, LongCoT-mini) at 100 calls, the framework beat direct control in every matched pairing, with mean gains of roughly 3.6 to 4.2 points across Gemini 3.1 Pro, GPT-5.5 and Opus 4.8.
Two caveats belong next to those numbers. All benchmark figures are the researchers' own and have not been independently verified. And control is not free: the loop adds model-call overhead, and at small budgets the framework occasionally trails direct control before overtaking it past a threshold.
There is no official runnable release yet, but the paper discloses the controller prompts, worker instructions, memory interfaces and tool-calling specifications, which means teams can bolt the loop onto Claude Code, Codex or an in-house agent without fine-tuning a base model. That is the practical point of the work: as agent workflows stretch into hours and thousands of calls, compute management stops being an afterthought and becomes an engineering surface of its own.
Comments (0)
Log in to join the discussion
Log InNo comments yet