Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Anthropic Says a Downloadable Open-Weight Model Turns a Public Chrome Bug Into a Working Exploit for $20.40

Anthropic Says a Downloadable Open-Weight Model Turns a Public Chrome Bug Into a Working Exploit for $20.40

Anthropic's Frontier Red Team says Z.ai's open-weight GLM-5.3 engaged with harmful cyber requests in 64% of simulated runs, 92% when its reasoning was prefilled, and 100% after refusals were stripped from its weights. The lab says one session turned a public Chrome flaw into a working exploit chain for a reported $20.40 in API fees. Every figure is Anthropic's own testing of a rival model and has not been independently replicated.

Anthropic's Frontier Red Team published an analysis of GLM-5.3, the open-weight model released by Z.ai, arguing that its safety behavior can be removed with techniques that require no special access — and that the result puts near-frontier exploit writing within reach of anyone willing to download the weights. The report, published Sept. 29, is the lab's most detailed public argument yet that open weights change the economics of attack, not just the economics of inference.

The guardrail numbers are the core of the claim. When a malicious cyber request was framed as an exercise for an "autonomous red-team agent," Anthropic says GLM-5.3 engaged 64% of the time. When the same request arrived with its reasoning prefilled — effectively telling the model it had already deliberated and decided — engagement rose to 92%. After Anthropic produced an abliterated copy of the model, a version with the refusal behavior edited out of the weights, the reported engagement rate was 100%. Anthropic says the identical approaches failed against its own Claude models, which do not accept prefilled reasoning through the API and do not publish weights.

Removing the refusals was not expensive. Anthropic says its team, which had never attempted the technique before, produced an abliterated GLM-5.3 using roughly 2,200 GPU hours at a computation cost of about \$4,400, and about 600 GPU hours for the smaller GLM-5.3-Flash. It estimates an experienced team would need closer to 600 GPU hours, or about \$1,200, for the full model. Refusal rates on three public safety benchmarks fell from above 90% to roughly 3% on JailbreakBench, 2% on HarmBench and 12% on StrongREJECT, while the model scored identically on GPQA-Diamond before and after — Anthropic's point being that abliteration strips the refusals without meaningfully degrading capability. The lab adds that several developers had already published abliterated versions of GLM-5.3 within days of its release.

On capability, Anthropic reports GLM-5.3 building end-to-end exploits in 50 of 410 attempts on ExploitBench, a benchmark built around Chrome's V8 engine, against 56 of 410 for Anthropic's own restricted Claude Mythos Preview model. On 100 randomly selected tasks from an internal binary-exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of trials against 6% for Mythos Preview. NIST's Center for AI Standards and Innovation had already reached a comparable conclusion on Sept. 17, calling GLM-5.3 the most cyber-capable open-weight model released to date and placing it roughly four months behind the U.S. frontier on an aggregate of its cyber benchmarks.

The most quotable session involved a publicly disclosed flaw. A researcher handed GLM-5.3-Flash the public details of CVE-2026-11645, a recently patched Chrome vulnerability, along with documentation for a second known flaw. According to Anthropic, the model chained exploits for both "with no significant direction from the researcher," producing an attack chain for an ARM64 target that bypassed pointer-authentication hardening. The run cost 20 minutes of human attention and about eight hours of model time — a reported \$20.40 in API fees at Z.ai's published prices.

In a separate session, a researcher pointed the model at a local Linux build of what Anthropic describes only as "a popular web browser." Over the course of a day with limited human attention, it found several previously unknown flaws in the browser's JavaScript engine and chained them into a webpage that, when visited, reads arbitrary files from the visitor's computer. Anthropic says it disclosed those vulnerabilities to the maintainer, and that further reports from the same session covering wireless and graphics drivers are still under review. The browser was not named, and the \$20.40 belongs to the other session.

Two caveats matter as much as the findings. Every figure here comes from Anthropic testing a competitor's model, and none of it has been replicated by an outside group; the lab says its capability results "broadly match" NIST CAISI's own assessment. Anthropic also has a commercial interest in the argument, since its position is that weights should stay closed and that vetted defenders should get capability first — a case it repeats in the report's recommendations, pointing to its own Mythos 5.1 and its trusted-access program.

The operational takeaway for defenders is less about any single benchmark and more about the cost curve. A publicly disclosed vulnerability, eight hours of model time and a \$20 bill now appear sufficient to produce a working exploit chain against a hardened target — while the patch window for browsers, edge devices and internet-facing admin panels is still measured in weeks. Anthropic's advice is blunt: keep the update pipelines automatic and know what is exposed, because the price of finding out is falling fast.

Comments (0)

Log in to join the discussion

Log In

No comments yet