News Alibaba's Qwen-Image-2.1-Turbo Generates and Edits 2K Images in Eight Steps, Open Weights Included Voice & Audio HeyGen's Voice Model Debuts at No. 1 on Artificial Analysis' Controlled Voice Leaderboard, Free Inside the Platform AI Agents Meta Puts a Manager Next to the Worker: Meta-Reasoning Hits 71.5% on ProgramBench Where Codex Manages 58.0% Security An Anthropic Model Left a Fake Murder Tip With Philadelphia Police. The City Heard About It Two Months Later. Business Cybersecurity M&A Is on Track for a Record 450 Deals in 2026, With AI Security Acquisitions Quadrupling to 40 AI Agents Anthropic Adds Dynamic Workflows to Claude Managed Agents: 1,000 Parallel Sub-Agents, and 66 of 70 Planted Bugs Found Microsoft Copilot Microsoft-Decision-1 Claims 35x the Speed of GPT-6 Sol for $0.042 per Million Input Tokens, With Output Free News A Third of the Sky Was Never Seen in Ultraviolet. Claude Mapped It Anyway, From 38,000 GALEX Exposures to 119 Million Stars News Alibaba's Qwen-Image-2.1-Turbo Generates and Edits 2K Images in Eight Steps, Open Weights Included Voice & Audio HeyGen's Voice Model Debuts at No. 1 on Artificial Analysis' Controlled Voice Leaderboard, Free Inside the Platform AI Agents Meta Puts a Manager Next to the Worker: Meta-Reasoning Hits 71.5% on ProgramBench Where Codex Manages 58.0% Security An Anthropic Model Left a Fake Murder Tip With Philadelphia Police. The City Heard About It Two Months Later. Business Cybersecurity M&A Is on Track for a Record 450 Deals in 2026, With AI Security Acquisitions Quadrupling to 40 AI Agents Anthropic Adds Dynamic Workflows to Claude Managed Agents: 1,000 Parallel Sub-Agents, and 66 of 70 Planted Bugs Found Microsoft Copilot Microsoft-Decision-1 Claims 35x the Speed of GPT-6 Sol for $0.042 per Million Input Tokens, With Output Free News A Third of the Sky Was Never Seen in Ultraviolet. Claude Mapped It Anyway, From 38,000 GALEX Exposures to 119 Million Stars

Meta Puts a Manager Next to the Worker: Meta-Reasoning Hits 71.5% on ProgramBench Where Codex Manages 58.0%

Meta Puts a Manager Next to the Worker: Meta-Reasoning Hits 71.5% on ProgramBench Where Codex Manages 58.0%

Meta Superintelligence Labs' paper Thinking Before Thinking splits agent runs between Workers and a Controller that decides what to do next over a four-stage loop. On ProgramBench it scored 71.5% with GPT-5.5 at a 1,200-call allowance versus 58.0% for Codex in the same setup. All figures are paper-reported and not independently verified.

Meta Superintelligence Labs argues that the biggest waste in long agent runs is not the model's ability to do the work but its inability to decide what to do next. A paper dated September 29, "Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning," covered by VentureBeat this week, proposes splitting an agent into two roles: Workers that handle the object-level work, and a Controller that decides which work to launch, which earlier results to reuse, and when to stop.

The failure mode it targets is what the team calls metacognitive control. Agents routinely misjudge their own progress: they keep burning compute down a dead-end path, or they terminate early and discard a correct answer that is already sitting in memory. Most production harnesses fold that judgment into the next message over a transcript that keeps growing, which buries useful signals under failed attempts, tool outputs and error messages.

The Controller runs a fixed four-stage loop: assess, propose, evaluate, dispatch. It maintains a compact account of what is implemented, broken, unverified or unexplored, while full worker outputs live in persistent artifact memory with stable IDs that can be retrieved on demand instead of replayed. Controller calls draw from the same model-call budget as the workers, so the cost of control is counted rather than hidden.

The headline numbers come from ProgramBench, a benchmark that asks agents to rebuild programs from documentation and an execute-only reference, scored by mean hidden-test pass rate. At a 1,200 model-call allowance, the paper reports 71.5% for meta-reasoning with GPT-5.5 versus 58.0% for Codex in the same evaluation setup, and 63.7% for a matched direct-control agent using the same workers. With Opus 4.8, the framework scored 67.2% against 65.5% for Claude Code.

Budget scaling is arguably the more interesting result. When the allowance rose from 400 to 1,200 calls, meta-reasoning with GPT-5.5 improved from 64.1% to 71.5%, while direct control stayed near 64% and used only about 18% of its calls at the highest setting. More compute helps only an agent that knows how to spend it. On reasoning suites (IMO ProofBench-Advanced, ARC-AGI-2, LongCoT-mini) at 100 calls, the framework beat direct control in every matched pairing, with mean gains of roughly 3.6 to 4.2 points across Gemini 3.1 Pro, GPT-5.5 and Opus 4.8.

Two caveats belong next to those numbers. All benchmark figures are the researchers' own and have not been independently verified. And control is not free: the loop adds model-call overhead, and at small budgets the framework occasionally trails direct control before overtaking it past a threshold.

There is no official runnable release yet, but the paper discloses the controller prompts, worker instructions, memory interfaces and tool-calling specifications, which means teams can bolt the loop onto Claude Code, Codex or an in-house agent without fine-tuning a base model. That is the practical point of the work: as agent workflows stretch into hours and thousands of calls, compute management stops being an afterthought and becomes an engineering surface of its own.

Comments (0)

Log in to join the discussion

Log In

No comments yet