- Start
- Sep 30, 202690% CONFIDENCEfrom the source
Gemini 4 Argon Announced
- Google announced Gemini 4 Argon on 30 September 2026, calling it its "new frontier model" for real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense[1]
- It is rolling out first to a set of trusted cyber defenders through the Fairwind Program, and Google is taking part in the U.S. government's voluntary process for pre-release model access while it expands access gradually[1]
- Google will make it available to developers, enterprises and consumers "as soon as possible", starting with paid API customers and Google AI Ultra subscribers, so the broad release is still pending[1]
- Internally it already powers Google's workflows: agents freed over 300 TiB of memory across data centers, beat a quantum-subroutine baseline by 40%, and helped migrate C/C++ to Rust up to 800K+ lines for the Fuchsia OS Zircon kernel[1]
- In a demonstration for cyber defense it found a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models had missed[1]
Notable features
- Output limit expanded to 1 million tokens, up from 64K, so the model can reason through hundreds of thousands of tokens in one trajectory[1]
- Introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper; afterwards $4 and $20 per million[1]
- New state of the art on DeepSWE v1.1 at 77.9% (GPT-6 Astra 74.1%, Claude Opus 5.5 74.2%), and first on AutomationBench at 51.3% and on the Vals Index at 68.9%[1][2]
- LVBench long-video understanding 91.7%, Vals Finance Agent v2 65.4% and Harvey's Legal Agent Benchmark 19.6%[1][2]
- CWE-bench v1 vulnerability remediation 68.0%, tied for first; Google releases it without cyber guardrails to trusted defenders and its own teams[1][2]
- Safeguards: refusal of harmful requests with legitimate dual-use science preserved, top score on Gray Swan's indirect prompt injection benchmark, chain-of-thought and action monitoring that halts out-of-bounds execution, and hardened sandboxes[1][2]
Benchmarks[2]
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| Vals Index (Knowledge work) | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench (Knowledge work, Score) | 51.3% | 41.4% | 31.4% | 42.5% |
| Vals Finance Agent v2 (Knowledge work) | 65.4% | 53.5% | 58.9% | 58.6% |
| Harvey's Legal Agent Benchmark (Knowledge work) | 19.6% | 5.4% | 6.7% | 3.8% |
| DeepSWE v1.1 (Agentic coding) | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 (Agentic coding) | 55.0% | 65.5% | 56.3% | 62.3% |
| Vibe Code Bench (Agentic coding) | 91.9% | 89.6% | 90.3% | 90.3% |
| Terminal-bench 4.0 (Agentic coding) | 57.4% | 58.2% | 57.9% | 66.4% |
| PostTrainBench (ML engineering) | 45.3% | 44.3% | 40.2% | 49.3% |
| Terminal-Bench Science 0.1 (Science and math) | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench 2 (Science and math) | 88.8% | 85.4% | 68.6% | 73.1% |
| RiemannBench (Science and math) | 76.0% | 72.0% | 65.6% | 69.6% |
| GraphWalks (Long context, Up to 128k, BFS (F1)) | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks (Long context, 256k to 1M, BFS (F1)) | 84.2% | 71.8% | 65.0% | 66.8% |
| Agent's Last Exam (Computer use, Pass rate) | 39.5% | 34.2% | — | 38.2% |
| OSWorld-2.0 (Computer use, Offline subsetPartial score) | 69.2% | 72.6% | — | — |
| Chartography (Multimodal understanding) | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench (Multimodal understanding) | 91.7% | 87.5% | 79.7% | 83.7% |
| CWE-bench v1 (Cybersecurity) | 68.0% | 68.0% | 58.0% | 67.0% |
References 287% CONFIDENCEOverall confidence: 87%How well the pin's source and references back up its dates.Weighted average of how firmly 2 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argonblog.google· Posted Sep 30, 2026· Starts Sep 30, 2026 ✓· 50% of score
Google's post "Gemini 4 Argon: our next era of frontier intelligence" is dated 30 September 2026 and says "Today, we're announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders"; broad availability for developers, enterprises and consumers follows "as soon as possible" (see the estimated pin 4810).
- [2]85%Gemini - Google DeepMinddeepmind.google· Added Sep 30, 2026· Starts Sep 30, 2026· 50% of score
DeepMind's Gemini page now leads with Gemini 4 Argon and carries its benchmark table against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, plus the four safeguard areas Google is strengthening before broad availability.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.