- Start
- Mar 3, 202690% CONFIDENCEfrom the source
Gemini 3.1 Flash-Lite Released
- Google introduced Gemini 3.1 Flash-Lite on 3 March 2026, "our fastest and most cost-efficient Gemini 3 series model", built for high-volume developer workloads[1]
- It is faster than 2.5 Flash, with a 2.5X faster time to first answer token and a 45% increase in output speed at similar or better quality[1]
Notable features
- Priced at $0.25 per million input tokens and $1.50 per million output tokens[1]
- Suited to translation, content moderation, generating user interfaces and building simulations[1]
- In preview via the Gemini API in Google AI Studio and Vertex AI[1]
Benchmarks[3]
| Benchmark | Gemini 3.1 Flash-Lite | Gemini 2.5 Flash | Gemini 2.5 Flash-Lite | GPT-5 mini | Claude 4.5 Haiku | Grok 4.1 Fast |
|---|---|---|---|---|---|---|
| Input price ($/1M tokens, no caching, Lower is better) | $0.25 | $0.30 | $0.10 | $0.25 | $1.00 | $0.20 |
| Output price ($/1M tokens, Lower is better) | $1.50 | $2.50 | $0.40 | $2.00 | $5.00 | $0.50 |
| Output speed (Tokens / s) | 363 | 249 | 366 | 71 | 108 | 145 |
| Humanity’s Last Exam (Academic reasoning (full set, text + MM), No tools) | 16.0% | 11.0% | 6.9% | 16.7% | 9.7% | 17.6% |
| GPQA Diamond (Scientific knowledge, No tools) | 86.9% | 82.8% | 66.7% | 82.3% | 73.0% | 84.3% |
| MMMU-Pro (Multimodal understanding and reasoning, No tools) | 76.8% | 66.7% | 51.0% | 74.1% | 58.0% | 63.0% |
| CharXiv Reasoning (Information synthesis from complex charts) | 73.2% | 63.7% | 55.5% | 75.5% | 61.7% | 31.6% |
| Video-MMMU (Knowledge acquisition from videos) | 84.8% | 79.2% | 60.7% | 82.5% | — | 74.6% |
| SimpleQA Verified (Parametric knowledge) | 43.3% | 28.1% | 11.5% | 9.5% | 5.5% | 19.5% |
| FACTS Benchmark Suite (Factuality benchmark across grounding, parametric, search, and MM.) | 40.6% | 50.4% | 17.9% | 33.7% | 18.6% | 42.1% |
| MMMLU (Multilingual Q&A) | 88.9% | 86.6% | 84.5% | 84.9% | 83.0% | 86.8% |
| LiveCodeBench (Code generation (UI: 1/1/2025-5/1/2025)) | 72.0% | 62.6% | 34.3% | 80.4% | 53.2% | 76.5% |
| MRCR v2 (8-needle) (Long context performance, 128k (average)) | 60.1% | 54.3% | 30.6% | 52.5% | 35.3% | 54.6% |
| MRCR v2 (8-needle) (Long context performance, 1M (pointwise)) | 12.3% | 21.0% | 5.4% | — | — | 6.1% |
References 385% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 3 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-liteblog.google· Posted Sep 30, 2026· Starts Mar 3, 2026 ✓· 41% of score
Google's post "Gemini 3.1 Flash-Lite: Built for intelligence at scale" is dated 3 March 2026 and says it is rolling out in preview that day to developers through the Gemini API and to enterprises through Vertex AI.
- [2]75%Gemini (language model) - Wikipediaen.wikipedia.org· Added Sep 30, 2026· 41% of score
The Wikipedia article lists Gemini 3.1 Flash-Lite as released on 3 March 2026.
- [3]95%Gemini 3.1 Flash-Lite - Model Carddeepmind.google· Published Mar 3, 2026· 18% of score
DeepMind's model card gives the benchmark table against the model's predecessor and named rivals.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.