Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM

Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens

Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens

Cockroach Labs has published a five-month retrospective on MOLT Sinai, a coding-agent pipeline run against its MOLT migration tools since April 21. It merged 1,238 pull requests across more than a million lines of code with seven reverts and roughly $135,000 in Claude tokens, and only 92 of 1,322 issues needed a human.

Database company Cockroach Labs has published a five-month retrospective on an unusual experiment: a coding-agent pipeline organized like a teaching hospital, running against its MOLT database migration tools. Between April 21 and September 11 the pipeline, named MOLT Sinai, merged 1,238 pull requests spanning more than a million lines of code, reverted seven of them, and consumed roughly $135,000 in Claude tokens.

The metaphor is load-bearing. In a migration tool, one subtle bug can corrupt user data on its way into the database, so the team says it optimized not for hourly throughput but for a blunter question: how many of these PRs would we be embarrassed to have merged? The engineers wrote that they would rather have agents run ten times slower if it meant the output stayed shippable. In the hospital framing, GitHub issues are patients, merges are discharges, and each issue typically spends one to two days in the pipeline.

Every agent plays a named clinical role. A Triage Nurse assesses the incoming issue, a Fellow writes the diagnostic workup and treatment plan, a Review Attending hunts for faults in both, and a Discharge Nurse audits whether the review was actually done before anything merges. A human Chief takes escalations. Four more agents run outside the main flow: a Charge Nurse that revives stalled issues every 30 minutes, an Infection Control agent that locks the pipeline when main breaks, a Safety Department that writes weekly process reviews, and a Research Department that proposes new work. The whole thing runs on GitHub Actions, with labels tracking which stage each patient has reached.

A handful of rules carry most of the safety. A Fellow may not write code before a reviewed plan. Any workup exceeding 1,000 lines of code must be decomposed into sub-issues. An agent may not weaken a test to make it pass. When a Fellow gets stuck it writes an I-PASS handoff, a format borrowed from clinical shift changes, and the receiving agent must state what it understood before starting work. Approved human decisions go into an append-only precedent log that other agents cite instead of re-escalating.

The headline numbers: 1,322 issues discharged over five months, of which only 92 escalated to a human, at an average cost of about $84 per issue. The largest experiment was adding IBM Db2 support to MOLT. A planning agent split the parent issue into 15 sub-issues, two of which were decomposed again, and by close the pipeline had opened 32 sub-issues under it, merged 27 pull requests, sent work back 55 times and escalated nine issues, two to a human. The team notes the equivalent Oracle support work in 2024 took nine months and roughly $160,000 of engineering time; the Db2 token bill came to $4,172. The caveat matters: the agent-written Db2 code had not yet been fully verified by humans at publication.

The post is equally candid about what broke. One urgent one-line fix took 11 rework rounds over two days, driven by false claims in the pull-request description, test-quality objections and two rebases. That led to a circuit breaker that hands a task to an Attending after four consecutive new-finding rounds or six rounds of any kind. An audit found the Discharge Nurse was loading about 21,800 words of instructions, roughly 29,000 tokens, before it read a single pull request, and a July review of the 25 skill files flagged 23 percent of their 100,000 words as removable without changing a single gate. Nearly half of all issues, the team says, were filed by the hospital for itself.

The broader lesson the authors draw is that reliability came from the structure of the pipeline, not from the quality of any single model output, and that the gates are where the money goes. For teams weighing a similar structure, the writeup offers two concrete reference points: a measured method for keeping agent-written code acceptable on a correctness-critical codebase, and a measured bill for running it.

Comments (0)

Log in to join the discussion

Log In

No comments yet