- Start
- Aug 13, 202695% CONFIDENCEfrom deepmind.google
Gemini 3.7 Flash Released
- Google launched Gemini 3.7 Flash on 13 August 2026, just three weeks after 3.6 Flash, calling it a "direct result of developer feedback and algorithmic innovations"[1]
- It gains "strong gains" over 3.6 Flash in debugging and issue resolution, with DeepSWE v1.1 rising from 49.0% to 65.3% and FrontierCode 1.1 Main from 34.4% to 43.6%[1]
- In the Gemini app it rolled out first to Spark, for AI Pro and Ultra subscribers[1]
Notable features
- Introductory price of $0.75 per million input tokens and $3.75 per million output tokens until the end of 2026, half the price 3.6 Flash launched at[1][2]
- WebDev Arena Elo 1588 (3.6 Flash 1538); GDP.pdf document comprehension 34.0% vs 22.0%; AutomationBench 30.4% vs 17.0%[1]
- Terminal-bench 2.1 85.8%, OSWorld-2.0 47.9% and an Artificial Analysis Intelligence Index of 56[2]
- Updated safeguards against chemical, biological, radiological and nuclear misuse and cyber offense[1]
Benchmarks[2]
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | Claude Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |
|---|---|---|---|---|---|
| Input price ($/1M tokens) | $0.75* | $0.75* | $2.00 | $2.00 | $1.25 |
| Output price ($/1M tokens) | $3.75* | $3.75* | $10.00 | $12.00 | $4.25 |
| Artificial Analysis Intelligence Index (Composite model intelligence) | 56 | 52 | 55 | 57 | 57 |
| FrontierCode 1.1 Main (Production code quality, Score) | 43.6% | 34.4% | 42.7% | 41.3% | — |
| DeepSWE v1.1 (Long-horizon software engineering) | 65.3% | 48.6% | 53.8% | 69.6% | 54.9% |
| Code Arena (Web development, Elo) | 1588 | 1538 | 1541 | 1523 | 1535 |
| Terminal-bench 2.1 (Agentic terminal coding) | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 (General agent capabilities) | 14.9% | 5.4% | 14.6% | 20.8% | — |
| AutomationBench (Enterprise workflow automation, Private set) | 30.4% | 17.0% | 10.7% | 23.6% | — |
| GDPVal-AA v2 (Knowledge work, Elo) | 1525 | 1422 | 1598 | 1578 | 1628 |
| Harvey LAB-AA (Complex legal workflows) | 90.7% | 85.1% | 90.1% | 85.2% | — |
| GDP.pdf (Expert PDF document comprehension) | 34.0% | 22.0% | 28.0% | 24.7% | 16.0% |
| CharXiv Reasoning (Information synthesis from complex charts, No tools) | 84.5% | 85.2% | 77.0% | 85.9% | — |
| CharXiv Reasoning (Information synthesis from complex charts, With tools) | 88.7% | 89.4% | 88.3% | — | — |
| LVBench (Long video understanding) | 85.4% | 84.2% | 68.5% | 78.9% | — |
| GDM-MRCR v2 (8-needle) (Long context performance, 128k (average)) | 97.0% | 91.8% | 81.5% | 93.5% | — |
| OSWorld-2.0 (Agentic computer use) | 47.9% | 33.8% | — | 50.2% | — |
| Agent's Last Exam (Multimodal desktop and OS agent tasks, Pass rate) | 26.3% | 24.2% | 33.3% | 28.0% | — |
| HLE-Verified (Multidisciplinary expert reasoning) | 53.6% | 51.2% | 31.0% | 51.1% | — |
| BioMysteryBench (Bioinformatics research reasoning, Human solvable) | 87.1% | 80.6% | 87.5% | 83.8% | — |
| BioMysteryBench (Bioinformatics research reasoning, Human difficult) | 43.5% | 41.2% | 34.1% | 49.4% | — |
| LABBench2 (Biology real-world research tasks) | 82.1% | 76.1% | 80.1% | 81.2% | — |
References 292% CONFIDENCEOverall confidence: 92%How well the pin's source and references back up its dates.Weighted average of how firmly 2 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%9to5google.com/2026/08/13/gemini-3-7-flash-launch9to5google.com· Posted Sep 30, 2026· Starts Aug 13, 2026· 55% of score
9to5Google's launch report is dated 13 August 2026 and says Google "today announced Gemini 3.7 Flash"; DeepMind's[2] model card is published 13 August 2026.
- [2]95%Gemini 3.7 Flash - Model Carddeepmind.google· Published Aug 13, 2026· Starts Aug 13, 2026 ✓· 45% of score
DeepMind's model card, published 13 August 2026, lists the introductory prices and a benchmark table against 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.