Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM

Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones

Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones

A new TMLR paper from Sakana AI grades AI peer reviewers on whether they catch planted mistakes instead of how closely they mimic human reviews. Across 257 papers with 1,164 inserted contradictions, its Multi-Layered Review system caught 73.43 percent of core-claim errors at about $0.47 per review, against 14.81 percent for the best baseline.

Most benchmarks for AI peer reviewers ask a soft question: does the machine write something that sounds like a human review? A new paper from Sakana AI, accepted at the machine learning journal TMLR, argues that this measures style rather than substance, and swaps in a harder one: if you bury a factual contradiction in the core claim of a real paper, does the reviewer notice? The paper, Beyond Imitation, and its companion benchmark were published on October 9.

The Contradiction Benchmark was built from 257 Creative Commons-licensed papers drawn from ACL, AISTATS, CVPR and ICML 2025 plus NeurIPS 2024. Gemini 2.5 Pro mapped each paper into a knowledge graph linking claims, evidence, methods and implementation details. The distance of a node from the main claim sets severity, with distance zero meaning the core contribution itself was overturned. GPT-4.1 then rewrote one node at each distance into a contradiction, producing 1,164 test cases. An o3 judge scored every review ten times; it answered no on unmodified originals 99.9 percent of the time and recognized 86.8 percent of manually confirmed catches, which suggests the reported detection rates may be conservative.

The system under test, Multi-Layered Review, runs three agents on off-the-shelf Claude models with no GPU and no fine-tuning. An Appendix Agent on Claude Haiku 3.5 summarizes the experiments and implementation details. An optional Literature Review Agent on Claude Sonnet 4 places the paper in prior work via web search. The Review Agent, also on Sonnet 4, reads up to ten pages of main text in three passes modeled on the Three-Pass Approach: outline first, detailed read second, merged review third. The PDF goes straight in, so figures and equations survive. Cost lands at about $0.47 per review.

With four reviews combined, MLR caught 73.43 percent of distance-zero contradictions and 40.95 percent overall. A single review still caught 60.79 percent. The best baseline, AgentReview, managed 14.81 percent, with LLM-Review at 14.56 percent and AI Reviewer at 11.17 percent. An ablation separates the two factors: swapping GPT-4.1 for Claude Sonnet 4 inside the LLM-Review baseline lifted distance-zero detection from 14.56 to 35.40 percent, and the MLR design added roughly 25 more points on top. Both the model and the architecture move the number.

The gains shrink where it counts. On 211 real retracted papers from WithdrarXiv-Check, MLR matched the stated retraction reason in 26.07 percent of cases on a similar-match basis and 16.11 percent exactly, against 18.48 and 9.00 percent for the strongest baselines. On scores, MLR correlated with human reviewers at 0.586 on ICLR 2025 submissions, below the 0.742 that humans achieve with each other, and edged them 0.439 to 0.429 on ICML 2025. Humans weigh clarity and novelty; the system stresses validity and experiments. The paper also concedes the system remains vulnerable to hidden prompt injection.

The motivation is volume. ICLR submissions grew from 2,594 in 2020 to 19,525 in 2026, with roughly 62,000 abstract registrations estimated for 2027; at ICLR 2026, 18,054 reviewers produced 76,139 reviews. Sakana stated its position plainly on X: peer review needs support, not substitutes, and much of the existing development effort optimizes for imitation.

The useful takeaway for anyone building research agents is that detection and mimicry are different targets, and that a reviewer which reads a paper before judging it starts from a higher base than one that critiques in a single pass. But the prompt-injection result is the warning attached to the headline: a system that can be talked out of its own judgment by text inside the very document it is reviewing is not ready to referee anything unsupervised.

Comments (0)

Log in to join the discussion

Log In

No comments yet