Google has unveiled Gemini 4 Argon, its newest frontier AI model, with a focus on long, complex work in software engineering, enterprise research, legal and finance tasks, and defensive cybersecurity. Google DeepMind says Argon can autonomously find, validate and patch critical software vulnerabilities. The model has a 1 million-token output limit and is initially being offered to trusted cyber defenders through Google's Fairwind program rather than released broadly. Google published benchmark results showing Argon ahead of several rival models on many knowledge-work, coding, science and multimodal tests, although it also trails competitors on some tasks. Reuters, The Verge and TechCrunch all reported the September 30 launch. For now, the biggest practical detail is access: most developers cannot simply start using Argon today.
Table of contents
- What Google launched
- Why cybersecurity teams get it first
- What the benchmarks show
- The 1 million-token output limit
- Access and pricing
- What the launch means for AI users
- FAQ
What Google launched
Google DeepMind announced Gemini 4 Argon on September 30, describing it as a model built for complex, long-horizon workflows rather than short chat responses. Its target areas include software engineering, legal and financial knowledge work, science, multimodal analysis and cybersecurity defense.
The model is already being used inside Google, according to the company. Google says Argon has helped with software work such as debugging and code migration, while also handling large amounts of information in a single task.
The Verge reported that Google is starting the rollout with a limited group of trusted cyber defenders. TechCrunch likewise reported that Argon is entering the Fairwind program first, with wider access planned later.
This makes the launch different from a normal Gemini model release. Google is showing the capability publicly, but access is being phased in while the company gathers feedback and continues safety work.
Why cybersecurity teams get it first
Cybersecurity is one of the main reasons Google is treating Argon differently. Google says the model can find, validate and patch critical software vulnerabilities autonomously.
That capability has an obvious defensive use, but the same ability can create security problems if an advanced model is used to discover or exploit weaknesses without authorization. Google says Argon includes safeguards intended to resist prompt injection and prevent harmful use while still allowing legitimate security research.
The company is therefore giving early access to vetted defenders through Fairwind. The idea is to let security teams test the model against real problems before a much wider release.
Google's approach comes during a period when AI agents have increasingly been tested on tasks that involve real systems and external tools. Reuters reported on September 30 that the U.S. Federal Trade Commission had opened an investigation into AI companies including OpenAI and Anthropic over safety and consumer-risk questions. That does not establish any connection between the FTC inquiry and Argon's launch, but it shows the wider scrutiny facing increasingly autonomous AI systems.
What the benchmarks show
Google's published comparison includes Gemini 4 Argon alongside OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and Claude Opus 5.5.
On DeepSWE v1.1, a software-engineering benchmark, Argon scored 77.9%, compared with 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. On Vibe Code Bench, Google reports 91.9% for Argon, ahead of 90.3% for both Claude Fable 5.1 and Claude Opus 5.5.
The picture changes on other coding tasks. Argon scored 55.0% on FrontierSWE v2, while GPT-6 Astra scored 65.5% and Claude Opus 5.5 scored 62.3%. On Terminal-bench 4.0, Argon recorded 57.4%, below GPT-6 Astra at 58.2% and Claude Opus 5.5 at 66.4%.
Google also reports strong results in knowledge work. Argon scored 68.9% on the Vals Index, 51.3% on AutomationBench, 65.4% on Vals Finance Agent v2 and 19.6% on Harvey's Legal Agent Benchmark. Those numbers come from Google's published evaluation table, so they should be treated as vendor-reported results rather than independent confirmation.
The model also scored 91.7% on LVBench, a multimodal benchmark focused on long-video understanding, according to Google DeepMind. On CWE-bench v1, a cybersecurity benchmark, Argon and GPT-6 Astra both scored 68.0%.
These results show a mixed picture rather than a clean sweep. Argon leads several tests, while other models lead on terminal-heavy coding and some computer-use tasks. Independent testing will matter more once outside researchers can access the model at scale.
The 1 million-token output limit
One of Argon's clearest technical changes is its output limit. Google says the model can produce up to 1 million tokens in a single response.
That does not mean a user will normally receive a million tokens every time they ask a question. The value is in long-running tasks where an AI system has to keep working through a large problem without constantly stopping for another turn.
For developers, that could matter in large code migrations, lengthy research workflows and tasks involving large collections of documents. For enterprise users, the same capacity could be useful when an agent needs to maintain more state while completing a multi-step job.
The practical limit will still depend on the task, tools, latency and reliability of the model. A bigger output window alone does not guarantee that an agent will complete a complicated job correctly.
Access and pricing
Argon is not broadly available at launch. Google says trusted cyber defenders are receiving access first through Fairwind, followed by a wider rollout to API customers and Google AI Ultra subscribers.
The Indian Express reported introductory API pricing of $2 per million input tokens and $10 per million output tokens. Google has not given a firm date for full public availability in the launch announcement.
That staged rollout also means developers cannot yet independently reproduce all of Google's benchmark claims using the production model. As access expands, third-party testing should provide a clearer picture of how Argon performs on real applications rather than curated evaluations.
What the launch means for AI users
Gemini 4 Argon puts Google back into the current frontier-model conversation at a time when AI labs are releasing models at a rapid pace. OpenAI launched GPT-6.1 Sol at DevDay just a day before Google's announcement, while Anthropic has been pushing its own high-end models and agent capabilities.
For developers, the interesting part is less about a single benchmark number and more about where Argon is being aimed. Google is building it around long-running engineering and enterprise tasks, while putting cybersecurity near the center of the first release.
For businesses, the limited rollout means there is still a gap between Google's claims and what an ordinary team can test today. The model may perform extremely well on some workloads and less well on others. Google's own table already shows that variation.
For security teams, the early access is more significant. If Argon can reliably discover and validate real vulnerabilities while staying inside authorization boundaries, it could become a useful defensive tool. But that claim needs testing outside Google's own evaluation environment.
For context, our earlier coverage of OpenAI's GPT-6.1 Sol launch and Nvidia's OpenShell and Sentry security platform shows how quickly the industry is moving toward AI systems that can take actions rather than simply generate text.
The next question is straightforward: when Argon reaches a wider developer audience, does its real-world performance match the benchmark lead Google is reporting?
FAQ
What is Gemini 4 Argon?
Gemini 4 Argon is Google's latest frontier AI model, designed for complex software engineering, enterprise knowledge work, multimodal tasks and defensive cybersecurity.
Can I use Gemini 4 Argon now?
Not broadly. Google is first giving access to trusted cybersecurity defenders through its Fairwind program, with wider access planned for API customers and Google AI Ultra subscribers.
How large is Gemini 4 Argon's output limit?
Google says Argon supports up to 1 million output tokens, allowing it to handle much longer multi-step responses than models with smaller output limits.
Is Gemini 4 Argon better than GPT-6 Astra or Claude?
Google's published benchmarks show Argon leading on several tests, but not all of them. It trails rival models on some coding and computer-use evaluations, so there is no single benchmark result that establishes it as better for every task.
Sources
Sources checked for this report include Google DeepMind's Gemini 4 Argon announcement, Reuters' coverage of the launch and current AI safety scrutiny, The Verge's report on Argon's limited cybersecurity rollout, TechCrunch's launch report, and The Indian Express' report on access and pricing.
