- 开始
- 2025年8月5日可信度 90%来自来源
Claude Opus 4.1 发布
由原文自动翻译
原标题: Claude Opus 4.1 Released
- Anthropic 于 2025 年 8 月 5 日向付费 Claude 用户、Claude Code 以及其 API、Amazon Bedrock 和 Google Cloud 的 Vertex AI 发布了 Claude Opus 4.1,价格与 Opus 4 相同[1][2]
- 该模型发布当天正值 OpenAI 发布其自 2019 年以来首批开放权重推理模型;Anthropic 表示计划“在未来几周对我们的模型做出更大幅度的改进”[1][6]
- GitHub 同日将其作为面向 Enterprise 和 Pro+ 套餐的公开预览版本纳入 Copilot[5]
- 与 Opus 4 一样,该模型按照 ASL-3 安全标准部署[3]
- 于 2026 年 8 月 5 日——恰好在发布一年后——从 Claude API 退役,由 Claude Opus 4.8 取代[4]
值得关注的功能
- 在“智能体任务、实际编程和推理方面”对 Claude Opus 4 进行了升级,深入研究和数据分析能力也有所提升,“尤其是在细节追踪和智能体搜索方面”[1]
- SWE-bench Verified 得分 74.5%,高于 Opus 4 的 72.5%;Terminal-Bench 得分 43.3%,高于 39.2%[1][6]
- Anthropic 表格中的其他成绩:GPQA Diamond 得分 80.9%,TAU-bench 零售场景得分 82.4%,MMMLU 得分 89.5%,AIME 2025 得分 78.0%[1]
- 客户反馈:GitHub 发现在多文件重构方面有显著提升,乐天(Rakuten)在大型代码库中获得了更精确的修复,“不会做出不必要的调整”,Windsurf 则观察到相较 Opus 4 有一个标准差的提升[1]
- 定价与 Opus 4 相同,每百万输入 token 15 美元、每百万输出 token 75 美元;API ID 为 claude-opus-4-1-20250805[1][6]
基准测试[1]
| 基准测试 | Claude Opus 4.1 | Claude Opus 4 | Claude Sonnet 4 | OpenAI o3 | Gemini 2.5 Pro |
|---|---|---|---|---|---|
| SWE-bench Verified¹(智能体编程) | 74.5% | 72.5% | 72.7% | 69.1% | 67.2% |
| Terminal-Bench²(智能体终端编程) | 43.3% | 39.2% | 35.5% | 30.2% | 25.3% |
| GPQA Diamond(研究生水平推理) | 80.9% | 79.6% | 75.4% | 83.3% | 86.4% |
| TAU-bench(智能体工具使用) | 零售 82.4%,航空 56.0% | 零售 81.4%,航空 59.6% | 零售 80.5%,航空 60.0% | 零售 70.4%,航空 52.0% | — |
| MMMLU³(多语言问答) | 89.5% | 88.8% | 86.5% | 88.8% | — |
| MMMU, validation(视觉推理) | 77.1% | 76.5% | 74.4% | 82.9% | 82% |
| AIME 2025⁴(高中数学竞赛) | 78.0% | 75.5% | 70.5% | 88.9% | 88% |
¹ Opus 4.1、Opus 4 和 Sonnet 4 配合 bash/编辑器工具运行 pass@1,10 次试验平均,单次尝试补丁,不使用测试时计算。² 默认智能体框架(Terminus 1),5 次试验平均。³ Claude 的成绩为 14 种非英语语言的平均值。⁴ 使用核采样(nucleus sampling),top_p 为 0.95。
参考资料 6可信度 89%总体可信度: 89%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 6 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/news/claude-opus-4-1anthropic.com· 发布于 2026年9月28日· 开始 2025年8月5日 ✓· 占评分 23%
Anthropic[3] 的发布文章标注日期为“2025 年 8 月 5 日”:“今天我们发布 Claude[2][4] Opus 4.1……现已面向付费 Claude 用户和 Claude Code 开放。它也已登陆我们的 API、Amazon Bedrock 和 Google Cloud 的 Vertex AI”;Claude 的发布说明和 GitHub[5] 的更新日志标注的日期相同。
- [2]90%Claude Platform release notes - Claude Platform Docsplatform.claude.com· 添加于 2026年9月28日· 占评分 23%
"August 5, 2025 We've launched Claude[4] Opus 4.1, an incremental update to Claude Opus 4", which does not allow both temperature and top_p to be set; later entries record structured outputs launching for Opus 4.1 (14 November 2025) and its retirement.
- [3]90%System Card Addendum: Claude Opus 4.1anthropic.com· 添加于 2026年9月28日· 占评分 23%
Anthropic's[1] system card addendum: "Like Claude[2][4] Opus 4, Claude Opus 4.1 is deployed under the AI Safety Level 3 (ASL-3) Standard"; it is not "notably more capable" than Opus 4 under the RSP, so new evaluations were voluntary.
- [4]90%Model deprecations - Claude Platform Docsplatform.claude.com· 添加于 2026年9月28日· 占评分 23%
Anthropic[1][3] notified developers on 5 June 2026 that Claude[2] Opus 4.1 (claude-opus-4-1-20250805) would be retired; it was retired on 5 August 2026, one year after release, with claude-opus-4-8 as the replacement.
- [5]85%Anthropic Claude Opus 4.1 is now in public preview in GitHub Copilotgithub.blog· 发表于 2025年8月5日· 开始 2025年8月5日· 占评分 5%
GitHub's changelog, "Release August 5, 2025": Claude[2][4] Opus 4.1, "the successor to Claude Opus 4", is available in GitHub Copilot Chat for Copilot Enterprise and Pro+ plans on github.com, Visual Studio Code and GitHub Mobile.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。