Google has announced Gemini 4 Argon, a new flagship AI model with an output limit of 1 million tokens and an initial rollout restricted to vetted cybersecurity defenders. The company describes it as its most capable model yet, with improvements in software engineering, enterprise tasks and vulnerability detection, although broader access has no announced timetable.
Unveiled on September 30 by Koray Kavukcuoglu, Google DeepMind’s senior vice president and Google’s chief AI architect, Argon is intended for complex tasks requiring sustained, multi-step work. Its launch combines company-reported benchmark gains with examples of internal deployments, while limiting early availability through Google’s Fairwind Program to selected security organizations and government partners.
Google says the model can identify, validate and patch critical software vulnerabilities autonomously. The company also says Argon discovered a vulnerability in healthcare software used by hospitals around the world that earlier frontier models had missed. Details about the affected software and the flaw were not provided in the supplied announcement account. Cloud security company Wiz is using the model through its Scan for Good initiative, which focuses on critical public infrastructure.
The restricted rollout reflects the dual-use nature of advanced cybersecurity tools: capabilities that help defenders locate weaknesses can also assist attackers. Google is providing approved defenders with access without its cyber-specific guardrails, while separately strengthening protections against prompt injection, misalignment and misuse. Removing cyber restrictions for vetted users is distinct from removing every safeguard.
In Google’s reported evaluations, Argon scored 77.9% on DeepSWE v1.1, a test of extended software engineering work involving real codebases, and 68.9% on the Vals Index, which covers enterprise tasks across finance, programming, law and tax. It also recorded 65.4% on Vals Finance Agent v2 and 51.3% on AutomationBench, an evaluation of end-to-end business execution. Its results were not uniformly ahead of competitors: on FrontierSWE v2, Argon’s 55.0% score trailed the reported results for GPT-6 Astra and Claude Opus 5.5.
The expanded output ceiling is another major technical change, rising from 64,000 tokens to 1 million. Output capacity determines how much a model can generate; it should not be confused with the input context window, which governs how much supplied material it can process. A larger generation allowance can support longer agent workflows, but does not by itself establish that a system will remain accurate or complete a task reliably.
Google’s internal examples include work on libgav1, its open-source video decoder. Argon agents reworked an existing Rust port, replacing 32,000 lines of SIMD code with safe Rust designed to allow automatic compiler vectorization. Google reports that the resulting decoder ran 2.7 times faster than the earlier Rust port while producing identical video output. In another deployment, agents analyzing data-center profiling information identified and applied optimizations that freed more than 300 TiB of memory.
Questions remain about how those results translate to everyday development. Bloomberg reported, citing employees with direct model access, that some internal testers found weaknesses in coding tasks despite strong benchmark scores. Google disputed that characterization. Benchmark outcomes and selected internal projects measure different aspects of performance, and neither establishes how the model will behave across all customer workloads.
Introductory pricing is set at $2 per million input tokens and $10 per million output tokens, with cached inputs costing $0.10 per million tokens. Standard prices will rise to $4 for input and $20 for output after an introductory period whose end date has not been specified. Google plans to extend access next to paid API customers and Google AI Ultra subscribers, but has not provided a release schedule.
Sources: BBN Times