Google today announced Gemini 4 Argon, a new frontier AI model that the company says leads rival systems on several enterprise-relevant benchmarks, including long-horizon software engineering, cybersecurity vulnerability remediation, business automation and economic-impact knowledge work.
The model is not yet broadly available. Google says Argon is beginning rollout to a set of trusted cyber defenders through its Fairwind Program introduced earlier this month, and that it is participating in the U.S. government’s voluntary pre-release model access process before a wider release.
Broad availability is planned “as soon as possible” for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers, according to an embargoed Google blog post reviewed by VentureBeat.
For enterprise buyers, the announcement matters less as a chatbot launch than as a signal that Google is trying to reassert itself in frontier AI after months of pressure from OpenAI and Anthropic.
The company is positioning Argon around three areas where enterprises are already spending: software development, professional knowledge work and cybersecurity operations.
Benchmark results: Argon leads or ties in the most categories, but not by a clean sweep

Google is not claiming that Gemini 4 Argon wins every benchmark. But across the benchmark table disclosed in the company’s embargoed materials, Argon posts the highest score, or ties for the highest score, in more categories than GPT-6 Astra or Claude Opus 5.5.
Across the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one. GPT-6 Astra leads outright on three and ties Argon on one. Claude Opus 5.5 leads outright on two. That makes Argon the leading frontier model by total number of benchmark leads or top scores in Google’s comparison set, even though the results show a still-close race in which OpenAI and Anthropic retain advantages in several important technical categories.
Argon’s strongest margins come in enterprise and long-context workflows. On Harvey’s Legal Agent Benchmark, Argon scores 19.6%, far ahead of GPT-6 Astra at 5.4% and Claude Opus 5.5 at 3.8%, making it one of the model’s clearest category wins. On AutomationBench, Zapier’s benchmark for end-to-end business execution, Argon scores 51.3%, compared with 42.5% for Claude Opus 5.5 and 41.4% for GPT-6 Astra. On the longer GraphWalks evaluation, which tests graph traversal over larger contexts, Argon scores 84.2%, compared with 71.8% for GPT-6 Astra and 66.8% for Claude Opus 5.5.
The model also posts meaningful leads in finance, coding and multimodal understanding. On Vals Finance Agent v2, Argon scores 65.4%, ahead of Claude Opus 5.5 at 58.6% and GPT-6 Astra at 53.5%. On DeepSWE v1.1, which measures real-world long-horizon software engineering tasks, Argon scores 77.9%, compared with 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. On LVBench, a long-video understanding benchmark, Argon scores 91.7%, ahead of GPT-6 Astra at 87.5% and Claude Opus 5.5 at 83.7%.
Those wins support Google’s broader claim that Argon is strongest across the kinds of workflows enterprises are likely to evaluate first: legal and financial analysis, coding, business automation, long-context reasoning, multimodal understanding and defensive cybersecurity. Argon also ties GPT-6 Astra on CWE-bench v1, with both models scoring 68% on the vulnerability-remediation benchmark, while Claude Opus 5.5 scores 67%.
But the benchmark table also shows where the race remains unsettled. Argon’s largest deficits are against GPT-6 Astra in FrontierSWE v2 and Terminal-Bench Science 0.1. In both cases, Astra leads by 10.5 percentage points: 65.5% to 55.0% on FrontierSWE v2, and 68.1% to 57.6% on Terminal-Bench Science 0.1. Claude Opus 5.5’s largest lead over Argon comes on Terminal-bench 4.0, where Opus scores 66.4% compared with Argon’s 57.4%, a 9-point gap. Opus also leads PostTrainBench, scoring 49.3% versus Argon’s 45.3%.
The result is a more nuanced claim than “best model across the board.” Argon appears to have the broadest top-score profile among the three frontier models in Google’s disclosed comparison, but GPT-6 Astra remains ahead in several software, science-terminal and computer-use tasks, while Claude Opus 5.5 remains ahead in terminal-agent and post-training workflows. For enterprise buyers, that means model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count.
For developers and CIOs, the practical takeaway is that Argon does not need a clean sweep to shift the competitive picture. By leading or tying in 13 of 18 disclosed categories, it gives Google the highest number of top scores across the comparison set. But the gaps where it trails — especially FrontierSWE v2, Terminal-Bench Science 0.1 and Terminal-bench 4.0 — show that OpenAI and Anthropic still have distinct areas of technical strength. The frontier race remains close, but Google can now argue that Argon leads it on breadth.
Argon also expands Google’s output ceiling. The company says the model supports an industry-leading 1 million output tokens, up from a previous 64,000-token limit. That is a notable distinction for agentic software engineering, audit, migration and legal-review workloads where the value of a model often depends on how long it can sustain a chain of work before handing control back to a human.
Google is already making heavy use of Argon internally
Inside Google, the company says Argon is already being used by thousands of employees for specialized coding tasks, deeper research and writing.
Google cites several internal examples: Argon helped quantum computing researchers optimize spacetime resources for bottlenecked subroutines, reportedly beating a published baseline by 40% in minutes; Argon agents analyzed fleet-wide profiling telemetry and identified memory optimizations expected to free more than 300 TiB across Google data centers once rolled out, with estimated total savings of 500 TiB to 1 PiB; and Argon agents are being used to migrate C/C++ codebases to Rust, including core libraries and the Fuchsia OS Zircon kernel.
The most concrete engineering example is libgav1, Google’s open-source video decoder.
According to Google’s blog post, Argon agents took an existing Rust port and replaced 32,000 lines of SIMD code by running profile-guided experiments, studying compiler output and producing safe Rust that the compiler could vectorize automatically. Google says the resulting memory-safe decoder runs 2.7 times faster than the Rust port while preserving identical video output.
Cybersecurity is the other major pillar. Google says Argon can autonomously find, validate and patch critical software vulnerabilities. For trusted defenders and Google’s own internal teams, the company says it will release Argon without cyber guardrails so those users can access its full defensive capabilities.
The company says Wiz is already using Argon through its Scan for Good initiative, which is focused on protecting critical public infrastructure, and that the model uncovered a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide.
Google is coupling that cyber-defense rollout with a safety message. Before broad availability, the company says it is strengthening safeguards against cyber and CBRN misuse, indirect prompt-injection attacks, model misalignment and insecure agent environments.
On the Gray Swan indirect prompt-injection benchmark shown in the blog, Gemini 4 Argon posts a 0.7% attack success rate, compared with 1.0% for Claude Opus 5.5 and Claude Fable 5.1, 8.5% for GPT-6 Astra, 27.0% for GPT-6 Sol, 31.5% for GLM 5.3 and 51.8% for Grok 4.8.
Pricing and availability: aggressive introductory pricing, but limited access
Google is pairing Argon’s benchmark claims with aggressive introductory API pricing. The model will launch at $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at a 95% discount to the input-token rate. That puts Argon’s introductory cached-input price at $0.10 per million tokens.
After the introductory period, Google says Argon will move to $4 per million input tokens and $20 per million output tokens.
Assuming the same 95% cached-input discount applies after the introductory period, cached input would rise to $0.20 per million tokens.
That pricing gives Google two different competitive stories. During the introductory period, Argon is priced at one-fifth of GPT-6 Astra’s listed API price of $10 per million input tokens and $50 per million output tokens, according to OpenAI’s API pricing page. It is also half the price of Claude Opus 5.5, which Anthropic lists at $4 per million input tokens and $20 per million output tokens.
After the introductory period, Argon’s standard price matches Claude Opus 5.5’s base API pricing at $4 input / $20 output per million tokens, while remaining below GPT-6 Astra’s $10 input / $50 output pricing. Anthropic also lists Claude Opus 5.5 cache reads at $0.20 per million tokens and cache writes at $5 per million tokens; Google’s stated 95% cached-input discount would make Argon cheaper on cached input during the introductory period and comparable after the introductory period, based on the disclosed rates.
The caveat is access. Argon is not launching immediately as a broadly available developer model. Google says it is first rolling out to trusted cyber defenders through its Fairwind Program while the company participates in the U.S. government’s voluntary pre-release model access process. Broader availability is planned for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers.
It is positioning Argon as both a benchmark leader by breadth and, at least during the introductory period, a cheaper frontier model than OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5.
The practical limitation is that most customers will still need to wait for broader API access before they can test whether Argon’s benchmark advantages translate into lower production costs on real coding, legal, financial and cybersecurity workloads.
Did Google successfully catch up after being behind for months?
The timing is important. Google has been under pressure to show that it can compete at the top of the model market, not only in distribution, infrastructure and cheaper Flash models. Last week, The Verge reported that Google’s new DeepMind chief, Koray Kavukcuoglu, said Gemini 4 was in refinement and could arrive well before year-end. The report noted that Google had not released a new flagship since the Gemini 3 series in November 2025, while OpenAI and Anthropic had advanced with GPT-6 and newer Claude models.
That gap followed months of public scrutiny. In July, Reuters reported that Alphabet faced investor concerns over a delayed Gemini 3.5 Pro release, rising AI infrastructure spending and departures from its AI teams. In August, Axios reported a major AI leadership shuffle: Demis Hassabis stepped down as CEO of Google DeepMind to become chairman and chief scientist at Alphabet; Jeff Dean left his chief scientist role to start a company with other AI researchers; and Kavukcuoglu moved into the top DeepMind role reporting to Sundar Pichai. Axios characterized the changes as Google’s biggest AI leadership shakeup since OpenAI’s 2023 upheaval, while noting that Google had not tied the moves to model delays.
External reporting has also pointed to deeper tensions. MarketWatch reported that Google’s AI organization had been affected by talent losses, delayed model releases and internal philosophical rifts over whether to prioritize scientific research, frontier capability or commercial products. Google’s announcement of Argon is therefore arriving not just as a model release, but as a response to a narrative that the company’s AI bench was deep but its frontier product cadence had slipped.
Argon gives Google a more credible answer to that criticism, at least on paper. It leads or ties rival models on several benchmarks that map closely to enterprise deployment: DeepSWE for long-horizon coding, CWE-bench for vulnerability remediation, Vals Index for high-value knowledge work, AutomationBench for business execution and Gray Swan IPI for prompt-injection robustness. It also gives Google a differentiated rollout strategy by starting with cyber defenders while it completes safety work and government pre-release access.
The enterprise adoption question
The remaining test is commercialization. Enterprises will want to know Argon’s pricing, rate limits, data-governance terms, deployment surfaces and integration path across Gemini API, Vertex AI, Google Cloud, Workspace and developer tools. Google has strong distribution advantages, but those advantages matter only if Argon’s benchmark performance translates into reliable production behavior.
For now, Gemini 4 Argon is Google’s clearest attempt in months to re-enter the frontier-model conversation on enterprise terms: less about consumer chatbot personality, and more about whether an AI agent can safely rewrite code, audit systems, analyze regulated knowledge work and defend software before attackers exploit it.
