Database company Cockroach Labs has published a five-month retrospective on an unusual experiment: a coding-agent pipeline organized like a teaching hospital, running against its MOLT database migration tools. Between April 21 and September 11 the pipeline, named MOLT Sinai, merged 1,238 pull requests spanning more than a million lines of code, reverted seven of them, and consumed roughly $135,000 in Claude tokens.
The metaphor is load-bearing. In a migration tool, one subtle bug can corrupt user data on its way into the database, so the team says it optimized not for hourly throughput but for a blunter question: how many of these PRs would we be embarrassed to have merged? The engineers wrote that they would rather have agents run ten times slower if it meant the output stayed shippable. In the hospital framing, GitHub issues are patients, merges are discharges, and each issue typically spends one to two days in the pipeline.
Every agent plays a named clinical role. A Triage Nurse assesses the incoming issue, a Fellow writes the diagnostic workup and treatment plan, a Review Attending hunts for faults in both, and a Discharge Nurse audits whether the review was actually done before anything merges. A human Chief takes escalations. Four more agents run outside the main flow: a Charge Nurse that revives stalled issues every 30 minutes, an Infection Control agent that locks the pipeline when main breaks, a Safety Department that writes weekly process reviews, and a Research Department that proposes new work. The whole thing runs on GitHub Actions, with labels tracking which stage each patient has reached.
A handful of rules carry most of the safety. A Fellow may not write code before a reviewed plan. Any workup exceeding 1,000 lines of code must be decomposed into sub-issues. An agent may not weaken a test to make it pass. When a Fellow gets stuck it writes an I-PASS handoff, a format borrowed from clinical shift changes, and the receiving agent must state what it understood before starting work. Approved human decisions go into an append-only precedent log that other agents cite instead of re-escalating.
The headline numbers: 1,322 issues discharged over five months, of which only 92 escalated to a human, at an average cost of about $84 per issue. The largest experiment was adding IBM Db2 support to MOLT. A planning agent split the parent issue into 15 sub-issues, two of which were decomposed again, and by close the pipeline had opened 32 sub-issues under it, merged 27 pull requests, sent work back 55 times and escalated nine issues, two to a human. The team notes the equivalent Oracle support work in 2024 took nine months and roughly $160,000 of engineering time; the Db2 token bill came to $4,172. The caveat matters: the agent-written Db2 code had not yet been fully verified by humans at publication.
The post is equally candid about what broke. One urgent one-line fix took 11 rework rounds over two days, driven by false claims in the pull-request description, test-quality objections and two rebases. That led to a circuit breaker that hands a task to an Attending after four consecutive new-finding rounds or six rounds of any kind. An audit found the Discharge Nurse was loading about 21,800 words of instructions, roughly 29,000 tokens, before it read a single pull request, and a July review of the 25 skill files flagged 23 percent of their 100,000 words as removable without changing a single gate. Nearly half of all issues, the team says, were filed by the hospital for itself.
The broader lesson the authors draw is that reliability came from the structure of the pipeline, not from the quality of any single model output, and that the gates are where the money goes. For teams weighing a similar structure, the writeup offers two concrete reference points: a measured method for keeping agent-written code acceptable on a correctness-critical codebase, and a measured bill for running it.
Comments (0)
Log in to join the discussion
Log InNo comments yet