- 开始
- 2026年9月10日可信度 90%来自来源
DeepSeek-V4.1-Flash 发布
由原文自动翻译
原标题: DeepSeek-V4.1-Flash Released
- DeepSeek 于 2026 年 9 月 10 日发布了 DeepSeek-V4.1-Flash,称其为“我们新架构系列中最小的模型,具备原生视觉理解能力”;该模型已以
deepseek-flash之名在 DeepSeek API 中上线,并接入 DeepSeek 的网页版和移动应用[1][5] - V4-Flash 和实验性的 V4-Flash-Vision-Exp 已被停用,其模型名称现均指向 V4.1-Flash;DeepSeek 正在“逐步淘汰 V4-Pro”,自 9 月 14 日 UTC 时间 04:00 起,所有
deepseek-v4-pro请求都由 V4.1-Flash 处理,并按 V4.1-Flash 的价格计费,直到 V4.1-Pro 发布为止[1][3] - 新的、更低的 API 价格已于 9 月 10 日 UTC 时间 04:00 生效;SiliconANGLE 计算得出,原本调用 V4-Pro 的用户在高峰时段每百万输出 token 的费用从 3.96 美元降至 1.20 美元,降幅约为 70%[1][5]
- 该模型的权重已以 MIT 许可证发布在 Hugging Face 上[4][5]
- 路透社指出,此次发布正值 DeepSeek 开始为在上海科创板(STAR Market)首次公开募股做准备之际[6]
值得关注的功能
- 该模型是一款拥有 5520 亿参数的混合专家模型,采用全新的因果编码器-解码器(Causal Encoder-Decoder)架构,“输入仅激活 80 亿参数,输出激活 160 亿参数”,基于 45 万亿个文本和图像 token 从零开始预训练[1][4]
- 其 KV 缓存每个 token 占用 890 字节,内存占用约为 V4-Flash 的四分之一,所需 SSD 存储空间仅为其八分之一;DeepSeek 指出,“缓存命中费用往往占智能体成本的很大比例”[1][4]
- 规格:1M token 上下文,最多 384K 输出 token,支持思考和非思考两种模式,推理强度设置范围为 1 到 100,原生支持图像输入、工具调用,并提供 Anthropic 格式的 API[2][4]
- 每百万 token 价格:非高峰时段输入 0.15 美元、输出 0.60 美元,高峰时段(工作日 UTC 时间 01:00-04:00 和 06:00-10:00)输入 0.30 美元、输出 1.20 美元,非高峰时段的缓存命中价格为 0.003 美元[2]
- 在最高推理强度下,DeepSeek 报告其在 Terminal-Bench 2.1 上得分 90.6,而 Claude Opus 5 为 89.1、GPT-5.6 Sol 为 88.8;在 DeepSWE v1.1 上得分 74.2%,对比 74.0% 和 73.0%(V4-Pro 为 62.7%);在 CyberGym 上得分 88.1;不过在 GPQA Diamond 上,两款美国模型仍领先(93.4 和 94.1,该模型为 90.9)[4][5]
参考资料 7可信度 85%总体可信度: 85%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 7 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%deepseek.com/en/news/deepseek-v4-1-flashdeepseek.com· 发布于 2026年9月28日· 开始 2026年9月10日 ✓· 占评分 15%
DeepSeek[2][3] 的公告日期为“2026 年 9 月 10 日”,文中写道:“V4.1-Flash 现已在 DeepSeek API 上线”;其 API 文档更新日志中列有“DeepSeek-V4.1-Flash 发布 2026/09/10”,路透社也报道 DeepSeek“今天(9 月 10 日)推出了 DeepSeek-V4.1-Flash”。
- [2]90%Models & Pricing | DeepSeek API Docsapi-docs.deepseek.com· 添加于 2026年9月28日· 占评分 15%
DeepSeek's[1][3] pricing page gives DeepSeek-V4.1-Flash a 1M context length and 384K maximum output, and per-million-token prices of $0.15 input (cache miss) and $0.60 output off-peak, $0.30 and $1.20 at peak, with peak hours 01:00-04:00 and 06:00-10:00 UTC on weekdays.
- [3]90%DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API Docsapi-docs.deepseek.com· 发表于 2026年9月10日· 开始 2026年9月10日· 占评分 14%
DeepSeek's[1][2] developer-docs copy of the announcement, listed as "DeepSeek-V4.1-Flash Release 2026/09/10": the model name deepseek-flash, the retirement of V4-Flash and V4-Flash-Vision-Exp, and "Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates".
- [4]90%deepseek-ai/DeepSeek-V4.1-Flash · Hugging Facehuggingface.co· 发表于 2026年9月10日· 占评分 14%
The official model card: 552B backbone parameters, a Causal Encoder-Decoder that activates 8B parameters per token in prefill and 16B in decode, a KV cache of 890 bytes per token, 45T pre-training tokens, the MIT license, and the benchmark table against Opus 5.0, GPT-5.6 Sol, Kimi K3 and GLM-5.3.[1]
- [5]85%DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro - SiliconANGLEsiliconangle.com· 发表于 2026年9月10日· 开始 2026年9月10日· 占评分 14%
SiliconANGLE's same-day report ("today released DeepSeek[1][2][3]-V4.1-Flash") works out that rerouted V4-Pro callers pay $1.20 instead of $3.96 per million output tokens at peak, "a cut of roughly 70%", and notes Anthropic named DeepSeek in a distillation report the same day.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。