- Start
- Feb 19, 202690% CONFIDENCEfrom the source
Gemini 3.1 Pro Released
- Google released Gemini 3.1 Pro in preview on 19 February 2026, three months after Gemini 3 Pro, describing it as the upgraded core intelligence behind recent science and research breakthroughs[1]
- It is "a smarter, more capable baseline for complex problem-solving", with a verified 77.1% on ARC-AGI-2, which tests entirely new logic patterns[1]
- Google kept it in preview to validate the updates and improve agentic workflows before general availability "soon"[1]
Notable features
- Available in the Gemini app with higher limits for Google AI Pro and Ultra plans, and in NotebookLM for Pro and Ultra users only[1]
- Developers and enterprises get it in preview through the Gemini API in AI Studio, Antigravity and Vertex AI, plus Gemini Enterprise[1]
- Google's demo has it code an interactive 3D starling murmuration that users steer with hand tracking and that plays a generative score[1]
Benchmarks[3]
| Benchmark | Gemini 3.1 Pro | Gemini 3 Pro | Sonnet 4.6 | Opus 4.6 | GPT-5.2 | GPT-5.3-Codex |
|---|---|---|---|---|---|---|
| Humanity's Last Exam (Academic reasoning (full set, text + MM), No tools) | 44.4% | 37.5% | 33.2% | 40.0% | 34.5% | — |
| Humanity's Last Exam (Academic reasoning (full set, text + MM), Search (blocklist) + Code) | 51.4% | 45.8% | 49.0% | 53.1% | 45.5% | — |
| ARC-AGI-2 (Abstract reasoning puzzles, ARC Prize Verified) | 77.1% | 31.1% | 58.3% | 68.8% | 52.9% | — |
| GPQA Diamond (Scientific knowledge, No tools) | 94.3% | 91.9% | 89.9% | 91.3% | 92.4% | — |
| Terminal-Bench 2.0 (Agentic terminal coding, Terminus-2 harness) | 68.5% | 56.9% | 59.1% | 65.4% | 54.0% | 64.7% |
| Terminal-Bench 2.0 (Agentic terminal coding, Other best self-reported harness) | — | — | — | — | 62.2% | 77.3% |
| SWE-Bench Verified (Agentic coding, Single attempt) | 80.6% | 76.2% | 79.6% | 80.8% | 80.0% | — |
| SWE-Bench Pro (Public) (Diverse agentic coding tasks, Single attempt) | 54.2% | 43.3% | — | — | 55.6% | 56.8% |
| LiveCodeBench Pro (Competitive coding problems from Codeforces, ICPC, and IOI, Elo) | 2887 | 2439 | — | — | 2393 | — |
| SciCode (Scientific research coding) | 59% | 56% | 47% | 52% | 52% | — |
| APEX-Agents (Long horizon professional tasks) | 33.5% | 18.4% | — | 29.8% | 23.0% | — |
| GDPval-AA Elo (Expert tasks) | 1317 | 1195 | 1633 | 1606 | 1462 | — |
| τ2-bench (Agentic and tool use, Retail) | 90.8% | 85.3% | 91.7% | 91.9% | 82.0% | — |
| τ2-bench (Agentic and tool use, Telecom) | 99.3% | 98.0% | 97.9% | 99.3% | 98.7% | — |
| MCP Atlas (Multi-step workflows using MCP) | 69.2% | 54.1% | 61.3% | 59.5% | 60.6% | — |
| BrowseComp (Agentic search, Search + Python + Browse) | 85.9% | 59.2% | 74.7% | 84.0% | 65.8% | — |
| MMMU-Pro (Multimodal understanding and reasoning, No tools) | 80.5% | 81.0% | 74.5% | 73.9% | 79.5% | — |
| MMMLU (Multilingual Q&A) | 92.6% | 91.8% | 89.3% | 91.1% | 89.6% | — |
| MRCR v2 (8-needle) (Long context performance, 128k (average)) | 84.9% | 77.0% | 84.9% | 84.0% | 83.8% | — |
| MRCR v2 (8-needle) (Long context performance, 1M (pointwise)) | 26.3% | 26.3% | — | — | — | — |
References 385% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 3 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-problog.google· Posted Sep 30, 2026· Starts Feb 19, 2026 ✓· 41% of score
Google's post "Gemini 3.1 Pro: A smarter model for your most complex tasks" is dated 19 February 2026 and says it is rolling out that day to developers, enterprises and consumers.
- [2]75%Gemini (language model) - Wikipediaen.wikipedia.org· Added Sep 30, 2026· 41% of score
The Wikipedia article lists Gemini 3.1 Pro as released in preview on 19 February 2026.
- [3]95%Gemini 3.1 Pro - Model Carddeepmind.google· Published Feb 19, 2026· 17% of score
DeepMind's model card gives the benchmark table against the model's predecessor and named rivals.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.