- Start
- Sep 10, 202690% CONFIDENCEfrom the source
DeepSeek-V4.1-Flash Released
- DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026, "the smallest model in our new architecture family, with native visual understanding", live on the DeepSeek API as
deepseek-flashand in DeepSeek's web and mobile apps[1][5] - V4-Flash and the experimental V4-Flash-Vision-Exp are retired and their model names route to V4.1-Flash; DeepSeek is "phasing out V4-Pro", and from 04:00 UTC on 14 September every
deepseek-v4-prorequest is answered by V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches[1][3] - New, lower API prices took effect at 04:00 UTC on 10 September; SiliconANGLE works out that a V4-Pro caller now pays $1.20 instead of $3.96 per million output tokens at peak, a cut of roughly 70%[1][5]
- The weights are on Hugging Face under the MIT license[4][5]
- Reuters notes the release came as DeepSeek starts preparing for an initial public offering on Shanghai's STAR Market[6]
Notable features
- A 552B-parameter mixture-of-experts model with a new Causal Encoder-Decoder architecture that activates "just 8B active parameters for input, 16B for output", pre-trained from scratch on 45T tokens of text and images[1][4]
- Its KV cache is 890 bytes per token, about 1/4 of V4-Flash's in memory, and needs 1/8 of the SSD storage; DeepSeek notes that "cache-hit charges often account for a large share of agent costs"[1][4]
- Specifications: 1M-token context, up to 384K output tokens, thinking and non-thinking modes, a reasoning-effort setting from 1 to 100, native image input, tool calls and an Anthropic-format API[2][4]
- Prices per million tokens: $0.15 input and $0.60 output off-peak, $0.30 and $1.20 at peak (weekdays 01:00-04:00 and 06:00-10:00 UTC), with cache hits at $0.003 off-peak[2]
- At maximum effort DeepSeek reports 90.6 on Terminal-Bench 2.1 against Claude Opus 5's 89.1 and GPT-5.6 Sol's 88.8, 74.2% on DeepSWE v1.1 against 74.0 and 73.0 (V4-Pro 62.7), and 88.1 on CyberGym; both US models still lead on GPQA Diamond (93.4 and 94.1 against 90.9)[4][5]
References 785% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 7 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%deepseek.com/en/news/deepseek-v4-1-flashdeepseek.com· Posted Sep 28, 2026· Starts Sep 10, 2026 ✓· 15% of score
DeepSeek's[2][3] announcement is dated "September 10, 2026" and says "V4.1-Flash is now live on the DeepSeek API"; its API-docs changelog lists "DeepSeek-V4.1-Flash Release 2026/09/10", and Reuters reports DeepSeek "today (10 September) launched DeepSeek-V4.1-Flash".
- [2]90%Models & Pricing | DeepSeek API Docsapi-docs.deepseek.com· Added Sep 28, 2026· 15% of score
DeepSeek's[1][3] pricing page gives DeepSeek-V4.1-Flash a 1M context length and 384K maximum output, and per-million-token prices of $0.15 input (cache miss) and $0.60 output off-peak, $0.30 and $1.20 at peak, with peak hours 01:00-04:00 and 06:00-10:00 UTC on weekdays.
- [3]90%DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API Docsapi-docs.deepseek.com· Published Sep 10, 2026· Starts Sep 10, 2026· 14% of score
DeepSeek's[1][2] developer-docs copy of the announcement, listed as "DeepSeek-V4.1-Flash Release 2026/09/10": the model name deepseek-flash, the retirement of V4-Flash and V4-Flash-Vision-Exp, and "Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates".
- [4]90%deepseek-ai/DeepSeek-V4.1-Flash · Hugging Facehuggingface.co· Published Sep 10, 2026· 14% of score
The official model card: 552B backbone parameters, a Causal Encoder-Decoder that activates 8B parameters per token in prefill and 16B in decode, a KV cache of 890 bytes per token, 45T pre-training tokens, the MIT license, and the benchmark table against Opus 5.0, GPT-5.6 Sol, Kimi K3 and GLM-5.3.[1]
- [5]85%DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro - SiliconANGLEsiliconangle.com· Published Sep 10, 2026· Starts Sep 10, 2026· 14% of score
SiliconANGLE's same-day report ("today released DeepSeek[1][2][3]-V4.1-Flash") works out that rerouted V4-Pro callers pay $1.20 instead of $3.96 per million output tokens at peak, "a cut of roughly 70%", and notes Anthropic named DeepSeek in a distillation report the same day.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.