Google on Wednesday revealed its next frontier model, Gemini 4 Argon — and then told almost everyone they cannot use it yet. Koray Kavukcuoglu, Google's chief AI architect and SVP at Google DeepMind, wrote in a blog post that Argon delivers "frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense." But instead of a public launch, the model is being released first to "a set of trusted cyber defenders," with Google saying it is "actively engaged in the U.S. government's voluntary process for pre-release model access" while it gradually expands availability.
The phased rollout runs through Google's Fairwind Program, which gives vetted defenders — and Google's own internal teams — a build of Argon without its cyber guardrails, so they can exercise the model's full defensive capabilities. Google frames the cautious approach as a necessity for models of this scale, and the move mirrors Anthropic's handling of Claude Mythos Preview, which has likewise stayed restricted to a small group of trusted organizations. Washington briefly forced Anthropic to suspend access to its publicly released models in June and has since set up a voluntary vetting process for the most powerful systems before release.
The benchmark sheet Google published is aimed squarely at rivals. On the software-engineering benchmark DeepSWE v1.1, Argon scores 77.9% — above GPT-6 Astra, Fable 5.1 and Opus 5.5, according to Google. It claims 68% on CWE-bench v1 (tied for first), 51.3% on AutomationBench (the top score on Zapier's benchmark) and 91.7% on LVBench for long-video understanding, which Google calls state of the art. All of these figures are company-published, and with the model locked away from the public, independent verification will have to wait.
Google also cited its own internal deployment as evidence. The company says Argon agents used fleet-wide telemetry data to free more than 300 TiB of data-center memory without added hardware, and have been migrating C/C++ codebases to Rust — including more than 800,000 lines in the Fuchsia OS Zircon kernel plus the re2 and libgav1 libraries, with the libgav1 video decoder now running 2.7 times faster. These are Google's own numbers, measured on Google's own infrastructure.
On safety, Google says Argon monitors the model's chain-of-thought and can halt a task mid-run if it strays out of bounds — a mechanism the company says flagged problems during training runs. The model is designed to refuse requests that could enable cyberattacks or the development of chemical, biological or nuclear weapons. As an early proof point, Google says security firm Wiz used Argon to uncover a critical vulnerability in software used by hospitals around the world that other frontier models had missed; Google did not name the flaw.
For developers, the headline specification is output: Argon can return up to 1 million tokens in a single response, up from 64,000 in previous Gemini models — enough, Google argues, to finish far larger tasks in one step. Chinese financial wire Cailianshe, quoting CEO Sundar Pichai, reported introductory pricing of $2 per million input tokens and $10 per million output tokens; Ars Technica noted that Google had not published full API pricing alongside the announcement, and the company gave no firm date for general availability, which will start with paid API users and Google AI Ultra subscribers.
Timing matters here. The reveal lands one day after OpenAI's DevDay, where the company launched its Dots agents and GPT-6.1 Sol and confirmed it would not ship the planned GPT-6.1 Astra over safety concerns. It also marks the first major frontier move for Kavukcuoglu since his August appointment atop DeepMind, and follows a summer in which Google promised Gemini 3.5 Pro for June but shipped only smaller Flash models.
The bigger signal is procedural, not technical. Between OpenAI cancelling a flagship release, Anthropic keeping its best model behind a vetting wall and now Google gating Argon to "trusted defenders," frontier-lab launches are turning into staged security clearances — with reasoning transparency recast as a safety mechanism and the public getting benchmarks and blog posts instead of a model to test.
Comments (0)
Log in to join the discussion
Log InNo comments yet