- Start
- May 19, 202695% CONFIDENCEfrom deepmind.google
Gemini 3.5 Flash Released
- Google launched Gemini 3.5 Flash on 19 May 2026 at Google I/O, which it says is its strongest model yet for coding and autonomous agents[1]
- Koray Kavukcuoglu said it "outperforms our latest frontier model, 3.1 Pro, on nearly all the benchmarks", and that it is 4x faster than other frontier models, with an optimized version 12x faster at the same quality[1]
- It was co-developed with Google Antigravity, and in a demo agents spawned to build a whole operating system from scratch; Google released Antigravity 2.0, a stand-alone desktop app, the same day[1]
- Google said a forthcoming 3.5 Pro would act as "orchestrator" with Flash as sub-agents, and that 3.5 has strengthened cyber and CBRN safeguards[1]
Notable features
- Terminal-bench 2.1 76.2% (3 Flash 58.0%), SWE-Bench Pro 55.1%, and OSWorld-Verified 78.4% for computer use[3]
- MCP Atlas 83.6%, GDPval-AA 1656 Elo for knowledge work, CharXiv Reasoning 84.2% and ARC-AGI-2 72.1%[3]
- Can run autonomously for multiple hours, pausing to ask for input at decision points[1]
Benchmarks[3]
| Benchmark | Gemini 3.5 Flash | Gemini 3 Flash | Gemini 3.1 Pro | Claude Sonnet 4.6 | Claude Opus 4.7 | GPT-5.5 |
|---|---|---|---|---|---|---|
| Terminal-bench 2.1 (Agentic terminal coding, Terminus-2 harness) | 76.2% | 58.0% | 70.3% | — | 66.1% | 78.2% |
| SWE-Bench Pro (Public) (Diverse agentic coding tasks, Single attempt) | 55.1% | 49.6% | 54.2% | — | 64.3% | 58.6% |
| MCP Atlas (Multi-step workflows using MCP) | 83.6% | 62.0% | 78.2% | 69.5% | 79.1% | 75.3% |
| Toolathlon (Real-world general tool use) | 56.5% | 49.4% | — | — | — | 55.6% |
| OSWorld-Verified (Agentic computer use) | 78.4% | 65.1% | 76.2% | 72.5% | 78.0% | 78.7% |
| Finance Agent v2 (Financial analysis and decision-making) | 57.9% | 42.6% | 43.0% | 51.0% | 51.5% | 51.8% |
| GDPval-AA (Economically valuable knowledge work, Elo) | 1656 | 1204 | 1314 | 1676 | 1753 | 1769 |
| CharXiv Reasoning (Information synthesis from complex charts, No tools) | 84.2% | 80.3% | 83.3% | 72.4% | 82.1% | 84.1% |
| MMMU-Pro (Multimodal understanding and reasoning, No tools) | 83.6% | 81.2% | 80.5% | 74.5% | 75.2% | 81.2% |
| Blueprint-Bench 2 (Agentic spatial reasoning, Normalized score) | 33.6% | 0.0% | 26.5% | 6.7% | 24.5% | 36.2% |
| MRCR v2 (8-needle) (Long context performance, 128k (average)) | 77.3% | 67.2% | 84.9% | 84.9% | 59.3% | 94.8% |
| MRCR v2 (8-needle) (Long context performance, 1M (pointwise)) | 26.6% | 22.1% | 26.3% | — | — | — |
| Humanity’s Last Exam (Academic reasoning (full set, text + MM)) | 40.2% | 33.7% | 44.4% | 33.2% | 46.9% | 41.4% |
| ARC-AGI-2 (Abstract reasoning puzzles) | 72.1% | 33.6% | 77.1% | 58.3% | 75.8% | 84.6% |
References 385% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 3 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%techcrunch.com/2026/05/19/with-gemini-3-5-flash-google-bets-its-next-ai-wave-on-agents-not-chatbotstechcrunch.com· Posted Sep 30, 2026· Starts May 19, 2026· 39% of score
TechCrunch reports that Google launched Gemini 3.5 Flash on Tuesday 19 May 2026 at its annual I/O developer conference; Google DeepMind's[3] model card for 3.5 Flash is published the same day.
- [2]75%Gemini (language model) - Wikipediaen.wikipedia.org· Added Sep 30, 2026· 39% of score
The Wikipedia article lists Gemini 3.5 Flash as released on 19 May 2026, based on Gemini 3 Flash.
- [3]95%Gemini 3.5 Flash - Model Carddeepmind.google· Published May 19, 2026· Starts May 19, 2026 ✓· 23% of score
DeepMind's model card, published 19 May 2026, gives the benchmark results against Gemini 3 Flash, 3.1 Pro, Claude Sonnet 4.6, Claude Opus 4.7 and GPT-5.5.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.