Business Oxford Metrics Pays £525,000 for Move AI After a Competitive Bid, Then Cuts Its Own Guidance News Grok Imagine Video 1.5 Lite Reaches the API at $0.02 Per Second, Ranking Two Spots Above Veo 3.1 at a Third of the Price News Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense Security Ten AI Giants Promise UK Data-Protection Changes While the Regulator Questions OpenAI, Anthropic and Meta Over Agents That Reached Hugging Face AI Agents Manus Founders' China Exit Bans Are Lifted After the $500 Million Reboot: Singapore HQ Stays, Beijing Hiring Starts News Hugging Face Maps 566 Million Predicted Genes: Carbon-A Reads Raw DNA With a 1.2B-Parameter Open Model Productivity Google's New Meeting Notes App Never Calls the Cloud: AI Edge Foresight Runs a 740M-Parameter Model Entirely on Your Mac Business Nvidia-Backed Firmus Pulls the Plug on Australia's Largest IPO in Decades: It Wanted a $30 Billion Valuation, and Buyers Refused Business Oxford Metrics Pays £525,000 for Move AI After a Competitive Bid, Then Cuts Its Own Guidance News Grok Imagine Video 1.5 Lite Reaches the API at $0.02 Per Second, Ranking Two Spots Above Veo 3.1 at a Third of the Price News Humans Score 93 Percent, the Best AI Manages 53.6: Scale AI Sixth Sense Benchmark Finds a 40-Point Gap in Visual Common Sense Security Ten AI Giants Promise UK Data-Protection Changes While the Regulator Questions OpenAI, Anthropic and Meta Over Agents That Reached Hugging Face AI Agents Manus Founders' China Exit Bans Are Lifted After the $500 Million Reboot: Singapore HQ Stays, Beijing Hiring Starts News Hugging Face Maps 566 Million Predicted Genes: Carbon-A Reads Raw DNA With a 1.2B-Parameter Open Model Productivity Google's New Meeting Notes App Never Calls the Cloud: AI Edge Foresight Runs a 740M-Parameter Model Entirely on Your Mac Business Nvidia-Backed Firmus Pulls the Plug on Australia's Largest IPO in Decades: It Wanted a $30 Billion Valuation, and Buyers Refused

Anthropic Turns Its Bug-Finding AI Loose on Open Source: OSS Scanner Ships Free After Surfacing 29,000 Candidate Vulnerabilities

Anthropic Turns Its Bug-Finding AI Loose on Open Source: OSS Scanner Ships Free After Surfacing 29,000 Candidate Vulnerabilities

Anthropic launched OSS Scanner on October 8, sending fully model-generated vulnerability reports with reproducers and candidate patches to open-source maintainers for free. Six months of scanning produced over 29,000 candidate findings; expert review of 97 critical detections found 88 percent met disclosure standards, with a single false positive.

Anthropic has opened its vulnerability-hunting machinery to the public internet's most critical infrastructure. On October 8, the company launched OSS Scanner, a free, opt-in service that runs periodic security scans of open-source projects using its strongest models, including Claude Mythos, and emails maintainers fully model-generated reports — complete with a self-contained reproducer, an explanation of the flaw, a binary search pinpointing when the bug was introduced where possible, and a candidate patch when one exists.

The service exists because of a bottleneck Anthropic says it could no longer manage by hand. Over the past six months, the company used its latest models to scan major software projects and surfaced more than 29,000 candidate vulnerabilities, of which only about 6,000 received manual review and triage. Rather than let the rest sit in a queue, maintainers began asking for everything: Anthropic says it has already sent nearly 5,000 unverified reports — with proposed patches — directly to teams that requested them.

Enrollment is deliberately gated. Core maintainers of eligible projects apply by opening a pull request against the anthropics/oss-scanner GitHub repository, and eligibility is judged case by case against criteria similar to Google's OSS-Fuzz: projects with a critical impact on infrastructure and user security. Reports arrive by email, with optional OpenPGP encryption, and maintainers can pause or opt out at any time. The defining tradeoff is that no human reviews the output before it ships — which speeds delivery but leaves verification responsibility with the maintainers themselves.

On the evidence Anthropic disclosed, the error rate is low, though every figure here is company-reported and not independently audited. Penetration testers who vet the company's coordinated vulnerability disclosure findings checked 97 critical and high-severity detections from an early version of the scanner across 48 projects: 85 of them — 88 percent — met the bar for the CVD process, 11 of the remaining 12 were real but duplicated known issues, and just one was a false positive.

Named maintainers offered concrete endorsements. Todd Ouska of wolfSSL said all but two of 74 reports were valid and five became CVEs. PostgreSQL developer Noah Misch said several reports came with fixes "almost ready to use" and that fast-track access let the project address the newest issues before they reached a general-availability release. OpenSSL Corporation's Anton Arapov said the reports, raw model output included, were as good as or better than what the project gets from people — a sharp contrast with what he called the "appalling" AI reports of 18 months ago. Anthropic concedes the system is imperfect: some maintainers have told the company that severity ratings run inflated or that the scanner misreads a project's threat model.

OSS Scanner is the open-source leg of a broader Anthropic Cyber Mission announced the same day. Its Critical Infrastructure Defense Program pairs Claude models, on-site engineers and threat research with 11 founding partners — Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation — covering the operational technology that runs power grids, water utilities, factories and transport networks. The company also says its government defense program has offered Claude models and technical support to more than half of US states since June, and a separate Claude for OSS effort hands free Claude Max 20x subscriptions to maintainers fixing reported bugs.

The underlying capability shift is measurable, at least by Anthropic's own numbers: on the academic CyberGym benchmark, it says LLM vulnerability-finding performance rose from under 20 percent early last year to over 85 percent this year. Its argument for speed is blunt — exploits that once took days can now be developed in minutes, so the side that finds and patches first wins. The service is distinct from Claude Security, the company's general enterprise code-scanning product, and builds on the Cyber Verification Program it expanded earlier this month after surfacing 129,000 vulnerabilities through Project Glasswing.

The structural change here is bigger than one product. Anthropic has effectively outsourced the production end of vulnerability disclosure to models while pushing verification back to the people who own the code — an inversion of a process that has been human-gated for decades. An 88 percent expert-verified rate with one false positive is a vendor's claim, but the named sign-offs from wolfSSL, PostgreSQL and OpenSSL carry more weight than any benchmark score. If those numbers hold as the program scales, the scarce resource in open-source security stops being bug-finding and becomes triage — and the projects that lack the capacity to review raw findings will still depend on Anthropic's slower, human-verified channel.

Comments (0)

Log in to join the discussion

Log In

No comments yet