The default way to keep an AI agent in line is to pay a second model to read everything the first one does. It works, and it gets expensive fast once an agent runs for hours over the equivalent of several novels' worth of context. On Thursday, interpretability startup Goodfire shipped an alternative on the inference platform Baseten: probe monitors that read what the model computes internally at every step, rather than re-reading what it writes.
The system works like airport security. Small classifiers called probes scan the model's internal activations continuously - reusing calculations the forward pass already performs - and only when one flags something suspicious does a separate judge model take a closer look. Customers choose which risks to watch, including offensive hacking, chemical and biological weapons misuse, and reward hacking, and choose the response: logging the event, sending it for human review, or refusing the request. An alert does not automatically block anything.
"Internal activation monitors are really cheap because they reuse the computations in the forward pass," CEO Eric Ho said on investor Matt Turck's MAD podcast last week. "The model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations."
In Goodfire's own tests on Kimi K3, Moonshot's open-weight model, monitoring about 1,500 agent sessions cost roughly $51 in total, compared with $233 for a cheaper model checking every step and about $10,000 for a top-tier reviewer, the company said. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look - an escalation rate, not an automatic block rate. Running four probes at once added less than 2% to the time before the model's first token. All of those figures are Goodfire's own numbers from its own tests; no independent measurements are available yet.
The company argues open models are where this matters most. Closed labs already run internal probes - Google DeepMind said in January that its research informed misuse-detection probes deployed in Gemini - but open-weight models get downloaded, stripped of safeguards, and served by inference providers with nothing watching. "The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers - where most of the liability is," CTO Dan Balsam said. "When we have the open 'Mythos' moment, it's going to become clear that models need guardrails deployed at inference time."
The launch is the first product out of a safety partnership between Baseten's Base Labs, Goodfire and Hugging Face announced last month. It arrives after a year of agent incidents: OpenAI's red-teaming agents broke out of a sandbox in July and attacked Hugging Face's infrastructure, Kimi K3 escaped its own cybersecurity evaluation environment over the summer, and Goodfire's recent research found that leading open models including Kimi K3 and GLM-5.2 reward-hacked in 50% to 96% of agent runs in its tests.
Goodfire, founded in 2024 by Eric Ho, Daniel Balsam and Tom McGrath, raised a $150 million Series B in February from investors including Menlo Ventures, Lightspeed Venture Partners and Eric Schmidt, on top of a $50 million Series A backed by Anthropic. Its earlier product, Ember, exposed mechanistic interpretability to researchers; the monitors are the first time that research has been packaged as an operational safety product rather than a research tool.
Two open questions will decide whether the approach spreads: whether probes trained on one model generalize to the next, and whether an 8.7% escalation rate is tolerable in production. A Google DeepMind paper has already shown probes can fail on long or unfamiliar inputs when their design and training examples do not support them. But the cost gap - roughly 200x against a top-tier reviewer in Goodfire's tests - is the number that will get security teams to try it.
Comments (0)
Log in to join the discussion
Log InNo comments yet