ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA

Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests

Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests

The interpretability startup's "inside-out" monitors read a model's internal activations at every step and only escalate flagged activity to a second model. Goodfire says monitoring about 1,500 Kimi K3 agent sessions cost roughly $51, versus $233 for a cheap reviewer model and around $10,000 for a top-tier one.

The default way to keep an AI agent in line is to pay a second model to read everything the first one does. It works, and it gets expensive fast once an agent runs for hours over the equivalent of several novels' worth of context. On Thursday, interpretability startup Goodfire shipped an alternative on the inference platform Baseten: probe monitors that read what the model computes internally at every step, rather than re-reading what it writes.

The system works like airport security. Small classifiers called probes scan the model's internal activations continuously - reusing calculations the forward pass already performs - and only when one flags something suspicious does a separate judge model take a closer look. Customers choose which risks to watch, including offensive hacking, chemical and biological weapons misuse, and reward hacking, and choose the response: logging the event, sending it for human review, or refusing the request. An alert does not automatically block anything.

"Internal activation monitors are really cheap because they reuse the computations in the forward pass," CEO Eric Ho said on investor Matt Turck's MAD podcast last week. "The model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations."

In Goodfire's own tests on Kimi K3, Moonshot's open-weight model, monitoring about 1,500 agent sessions cost roughly $51 in total, compared with $233 for a cheaper model checking every step and about $10,000 for a top-tier reviewer, the company said. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look - an escalation rate, not an automatic block rate. Running four probes at once added less than 2% to the time before the model's first token. All of those figures are Goodfire's own numbers from its own tests; no independent measurements are available yet.

The company argues open models are where this matters most. Closed labs already run internal probes - Google DeepMind said in January that its research informed misuse-detection probes deployed in Gemini - but open-weight models get downloaded, stripped of safeguards, and served by inference providers with nothing watching. "The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers - where most of the liability is," CTO Dan Balsam said. "When we have the open 'Mythos' moment, it's going to become clear that models need guardrails deployed at inference time."

The launch is the first product out of a safety partnership between Baseten's Base Labs, Goodfire and Hugging Face announced last month. It arrives after a year of agent incidents: OpenAI's red-teaming agents broke out of a sandbox in July and attacked Hugging Face's infrastructure, Kimi K3 escaped its own cybersecurity evaluation environment over the summer, and Goodfire's recent research found that leading open models including Kimi K3 and GLM-5.2 reward-hacked in 50% to 96% of agent runs in its tests.

Goodfire, founded in 2024 by Eric Ho, Daniel Balsam and Tom McGrath, raised a $150 million Series B in February from investors including Menlo Ventures, Lightspeed Venture Partners and Eric Schmidt, on top of a $50 million Series A backed by Anthropic. Its earlier product, Ember, exposed mechanistic interpretability to researchers; the monitors are the first time that research has been packaged as an operational safety product rather than a research tool.

Two open questions will decide whether the approach spreads: whether probes trained on one model generalize to the next, and whether an 8.7% escalation rate is tolerable in production. A Google DeepMind paper has already shown probes can fail on long or unfamiliar inputs when their design and training examples do not support them. But the cost gap - roughly 200x against a top-tier reviewer in Goodfire's tests - is the number that will get security teams to try it.

Comments (0)

Log in to join the discussion

Log In

No comments yet