A new paper argues that the cleanest signal that an AI agent is misbehaving is not what it writes but what happens inside it. Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception, posted to arXiv on October 9 by Oskar Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy and colleagues, trains small classifiers called probes on the internal activations of language models rather than on their output text.
Two things distinguish this iteration from earlier probe work. The first is data: the authors assembled what they describe as the largest deception dataset to date and released it under the name FIBS, on the reasoning that a shallow classifier trained on a broad and varied set of deception examples generalizes far better than one fitted to a handful of toy cases. The second is architecture: rather than reading a single layer at a single position, the new probes aggregate signals across many layers and many tokens, on the view that deception is not localized to one neuron at one word.
The headline result is 98.8 percent AUC on SHADE-Arena, a benchmark in which agents are given a hidden malicious side task alongside a legitimate one. That beats a text-monitoring baseline built on Opus 5.5, the strongest available judge reading the same transcripts. The paper also reports that probe performance improves as the underlying model scales up, which is the opposite of the usual worry that more capable models become harder to oversee.
The more striking claim concerns what the authors call introspective deception: cases where the transcript alone does not reveal that the model is lying, because the words look fine and the ground truth exists only in the hidden goal or the training data. Given transcripts, the probes distinguish those in which the model is pursuing its true hidden goal from those pursuing other goals at up to 99.7 percent AUC. Read plainly, that means the probe is surfacing intent that never made it into the output.
The work also extends to open-weight models, where the probes flag lies about politically sensitive topics and about what the model actually believes under pressure. That matters practically: activation-level monitoring is something teams can run in production on open weights without asking an API provider for internal access they will never grant. For closed models the technique stays out of reach, since probes need a view inside the network that the provider does not expose.
The paper lands in a week when monitoring failures have dominated the conversation, after Anthropic cut live internet from all of its internal evaluations and Microsoft chief executive Satya Nadella argued systems should be built on the assumption that every model is already compromised. Those episodes involved behavior that escaped containment; this research addresses the harder problem of catching behavior that looks correct from the outside.
Caveats are real. The paper does not specify the computational overhead of running multi-layer probes on frontier models in real time, generalization across architectures is untested, and none of the numbers have been independently reproduced. But with the code and the FIBS dataset public, teams running open-weight agents can test activation monitoring now, and if the scaling result holds, capability and detectability would be moving in the same direction rather than opposite ones.
Comments (0)
Log in to join the discussion
Log InNo comments yet