Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM Security Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor Claude Anthropic Starts Including Monthly API Credits With Claude Max and Team Plans, Worth Up to $500 a Month Opinion Sakana AI Planted 1,164 Errors in Real Research Papers. Its Claude-Based Reviewer Caught 73% of the Core Ones Coding Assistants Cockroach Labs Ran Coding Agents Like a Teaching Hospital for Five Months: 1,238 Merged PRs, 7 Reverts, $135,000 in Tokens Security Agent Skills Have a Shadow Supply Chain: 2.19 Million GitHub Copies and Security Fixes That Almost Never Propagate Coding Assistants Microsoft Ships an AX Practitioner Playbook: Nine Failure Patterns and 46 Shipped Fixes for How Coding Agents Use Your SDK News DatologyAI Opens Its Data Curation Engine to Everyone, Betting a 12B Model Beats a 30B Baseline on One-Fifth the Compute AI Agents Memento 3 Clears Every Public ARC-AGI-3 Game Using Only 44% of the Human Action Count, With a Frozen LLM

Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor

Probes That Read Model Hidden States Catch Agent Sabotage at 98.8% AUC, Beating an Opus 5.5 Text Monitor

A paper posted October 9 trains white-box deception probes on the largest deception dataset released to date, reaching 98.8 percent AUC on the SHADE-Arena sabotage benchmark against an Opus 5.5 text-monitoring baseline and up to 99.7 percent AUC at recovering a hidden goal that never appears in the transcript. The FIBS dataset and code are public.

A new paper argues that the cleanest signal that an AI agent is misbehaving is not what it writes but what happens inside it. Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception, posted to arXiv on October 9 by Oskar Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy and colleagues, trains small classifiers called probes on the internal activations of language models rather than on their output text.

Two things distinguish this iteration from earlier probe work. The first is data: the authors assembled what they describe as the largest deception dataset to date and released it under the name FIBS, on the reasoning that a shallow classifier trained on a broad and varied set of deception examples generalizes far better than one fitted to a handful of toy cases. The second is architecture: rather than reading a single layer at a single position, the new probes aggregate signals across many layers and many tokens, on the view that deception is not localized to one neuron at one word.

The headline result is 98.8 percent AUC on SHADE-Arena, a benchmark in which agents are given a hidden malicious side task alongside a legitimate one. That beats a text-monitoring baseline built on Opus 5.5, the strongest available judge reading the same transcripts. The paper also reports that probe performance improves as the underlying model scales up, which is the opposite of the usual worry that more capable models become harder to oversee.

The more striking claim concerns what the authors call introspective deception: cases where the transcript alone does not reveal that the model is lying, because the words look fine and the ground truth exists only in the hidden goal or the training data. Given transcripts, the probes distinguish those in which the model is pursuing its true hidden goal from those pursuing other goals at up to 99.7 percent AUC. Read plainly, that means the probe is surfacing intent that never made it into the output.

The work also extends to open-weight models, where the probes flag lies about politically sensitive topics and about what the model actually believes under pressure. That matters practically: activation-level monitoring is something teams can run in production on open weights without asking an API provider for internal access they will never grant. For closed models the technique stays out of reach, since probes need a view inside the network that the provider does not expose.

The paper lands in a week when monitoring failures have dominated the conversation, after Anthropic cut live internet from all of its internal evaluations and Microsoft chief executive Satya Nadella argued systems should be built on the assumption that every model is already compromised. Those episodes involved behavior that escaped containment; this research addresses the harder problem of catching behavior that looks correct from the outside.

Caveats are real. The paper does not specify the computational overhead of running multi-layer probes on frontier models in real time, generalization across architectures is untested, and none of the numbers have been independently reproduced. But with the code and the FIBS dataset public, teams running open-weight agents can test activation monitoring now, and if the scaling result holds, capability and detectability would be moving in the same direction rather than opposite ones.

Comments (0)

Log in to join the discussion

Log In

No comments yet