Google has released Mantis, an open-source toolkit under the Apache 2.0 license that turns an AI coding agent into a staged security reviewer. Rather than a scanner that fires at a repository and returns a pile of alerts, Mantis is a collection of slash-command skill directories that chain the agent through threat modeling, scanning, sandboxed reproduction, patching and verification - with the reproduction and re-attack stages treated as the actual trust boundary.
The problem it targets is specific and quantified by Google itself: naive AI code scanning yields true-positive rates below 7 percent, producing finding queues too noisy for security teams to act on. Mantis responds by refusing to accept a finding until it survives execution. Suspected issues are reproduced inside a gVisor container or a virtual machine with networking disabled, and any accepted patch is then re-attacked with a variant payload to confirm the fix is not simply bypassable.
The pipeline is organized in four phases. The learn phase builds context: /mantis-history mines version control for past security fixes, /mantis-summarize writes directory maps, /mantis-architecture assembles a linked Markdown knowledge base, /mantis-threat-model derives trust boundaries and /mantis-plan produces a review roadmap. The find phase sweeps files and filters: /mantis-researcher, /mantis-dedupe, /mantis-review and /mantis-critic collapse duplicates and drop issues that cannot occur in a release build. The prove phase executes payloads and assembles multi-step exploit chains, and the fix phase applies minimal patches, scores residual risk from 1 to 10 across impact, evidence and viability, and writes a human-readable report. A supervisor skill, /mantis-meta-agent, can drive the whole loop in a long-lived session.
Token economics are treated as a first-class constraint. Google says a hierarchical summary tree cuts token overhead by more than 85 percent compared with feeding a full repository to the model, which is what makes the pipeline economically viable to run repeatedly across large codebases. Mantis works with Gemini CLI, Antigravity CLI, the Google ADK or any comparable agent framework, and a newer skill called /mantis-advise inverts the flow - querying the accumulated threat model before new code is written so the same bug class does not land twice.
The caveats are unusually explicit for a corporate open-source release. Google positions Mantis as suitable for local and internal evaluation only, not production, and stresses that every generated finding or patch requires human security-expert verification. The project is not a supported Google product, and models remain non-deterministic enough to produce false positives, incorrect patches or unsafe actions - the documentation directs reproducers to isolated containers with networking disabled.
What makes Mantis worth attention is architectural honesty. Most agentic security tooling stops at generating findings and leaves orchestration to the model; Google publishes the inter-stage contracts so teams can wrap the skills in a deterministic harness instead of trusting an LLM to chain shell commands. As labs race to give coding agents the ability to find and fix vulnerabilities autonomously, the interesting question is not whether the model can write a patch - it is whether anyone can audit how the agent got there.
Comments (0)
Log in to join the discussion
Log InNo comments yet