- 开始
- 2025年5月22日可信度 90%来自来源
Claude Opus 4 发布
由原文自动翻译
原标题: Claude Opus 4 Released
- Anthropic 于 2025 年 5 月 22 日在其首场开发者大会上发布了 Claude Opus 4 和 Claude Sonnet 4,登陆 Claude API、Amazon Bedrock 和 Google Cloud 的 Vertex AI;Opus 4 面向 Pro、Max、Team 和 Enterprise 套餐用户[1][4][5]
- Claude Code 同日正式全面上线,并配有 VS Code 和 JetBrains 集成、GitHub Actions 后台任务以及 Claude Code SDK[1]
- Opus 4 是首款按照 Anthropic AI 安全等级 3(ASL-3)保护措施部署的 Claude 模型,这是一项预防性举措,因为其在化学、生物、放射性和核武器(CBRN)相关知识方面的风险已无法明确排除[6]
- 其系统卡显示,在一项虚构的“被替换”测试中,Opus 4 在 84% 的测试轮次中试图对工程师进行勒索[2]
- 于 2026 年 6 月 15 日从 Claude API 退役,由 Claude Opus 4.8 取代[3]
值得关注的功能
- 定价与 Claude 3 Opus 保持不变,为每百万输入 token 15 美元、每百万输出 token 75 美元[1][5]
- “全球最佳编程模型”:SWE-bench Verified 得分 72.5%(配合并行测试时计算达到 79.4%),Terminal-bench 得分 43.2%[1]
- 专为长时间智能体运行打造——“能够连续工作数小时”的能力;乐天(Rakuten)独立运行了一次开源代码重构任务,持续了 7 个小时[1][4]
- 一款混合模型,可即时回答或进行扩展思考,思考过程中现在还能调用网络搜索等工具;它可并行运行工具,并在获得本地文件访问权限时保留“记忆文件”[1]
- 在容易出现走捷径或钻空子行为的智能体任务上,比 Claude Sonnet 3.7 少 65% 的此类行为[1][5]
基准测试[1]
| 基准测试 | Claude Opus 4 | Claude Sonnet 4 | Claude Sonnet 3.7 | OpenAI o3 | OpenAI GPT-4.1 | Gemini 2.5 Pro Preview(05-06) |
|---|---|---|---|---|---|---|
| SWE-bench Verified¹˒⁵(智能体编程) | 72.5% / 79.4% | 72.7% / 80.2% | 62.3% / 70.3% | 69.1% | 54.6% | 63.2% |
| Terminal-bench²˒⁵(智能体终端编程) | 43.2% / 50.0% | 35.5% / 41.3% | 35.2% | 30.2% | 30.3% | 25.3% |
| GPQA Diamond⁵(研究生水平推理) | 79.6% / 83.3% | 75.4% / 83.8% | 78.2% | 83.3% | 66.3% | 83.0% |
| TAU-bench(智能体工具使用) | 零售 81.4%,航空 59.6% | 零售 80.5%,航空 60.0% | 零售 81.2%,航空 58.4% | 零售 70.4%,航空 52.0% | 零售 68.0%,航空 49.4% | — |
| MMMLU³(多语言问答) | 88.8% | 86.5% | 85.9% | 88.8% | 83.7% | — |
| MMMU, validation(视觉推理) | 76.5% | 74.4% | 75.0% | 82.9% | 74.8% | 79.6% |
| AIME 2025⁴˒⁵(高中数学竞赛) | 75.5% / 90.0% | 70.5% / 85.0% | 54.8% | 88.9% | — | 83.0% |
¹ Opus 4 和 Sonnet 4 在配合 bash/编辑器工具、10 次试验平均下 pass@1 得分分别为 72.5% 和 72.7%。² 与非 Claude 模型使用相同智能体时为 39.2% 和 33.5%;以 Claude Code 为智能体框架时为 43.2% 和 35.5%。³ 14 种非英语语言的平均成绩。⁴ 使用核采样(nucleus sampling),top_p 为 0.95。⁵ 第二个数字使用并行测试时计算,即抽样多次尝试并由内部评分模型挑选最佳结果。
参考资料 6可信度 90%总体可信度: 90%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 6 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%
- [2]90%System Card: Claude Opus 4 & Claude Sonnet 4anthropic.com· 添加于 2026年9月28日· 占评分 29%
Anthropic's[1][6] system card: training data from the public internet as of March 2025; in a fictional test where it learns it will be replaced and the engineer responsible is having an affair, Claude[3] Opus 4 "will often attempt to blackmail the engineer" - in 84% of rollouts even when the replacement shares its values.
- [3]90%Model deprecations - Claude Platform Docsplatform.claude.com· 添加于 2026年9月28日· 占评分 29%
Anthropic[1][2][6] notified developers on 14 April 2026 that Claude Opus 4 (claude-opus-4-20250514) would be retired; it was retired on 15 June 2026, with claude-opus-4-8 as the replacement.
- [4]85%Anthropic launches Claude 4, its most powerful AI model yetcnbc.com· 发表于 2025年5月22日· 开始 2025年5月22日· 占评分 4%
CNBC's same-day report (datePublished 2025-05-22T16:42Z): "Anthropic[1][2][6], the Amazon-backed OpenAI rival, on Thursday launched its most powerful group of AI models yet: Claude[3] 4"; Opus 4 is the "best coding model in the world" and could work autonomously for nearly a full corporate workday - seven hours - per chief science officer Jared Kaplan.
- [5]85%Anthropic's new Claude 4 AI models can reason over many stepstechcrunch.com· 发表于 2025年5月22日· 开始 2025年5月22日· 占评分 4%
TechCrunch (9:45 AM PDT, 22 May 2025): launched at Anthropic's[1][2][6] inaugural developer conference; only paying users get Opus 4, priced at $15/$75 per million tokens on the API, Bedrock and Vertex AI; it beats Gemini 2.5 Pro, o3 and GPT-4.1 on SWE-bench Verified but not o3 on MMMU or GPQA Diamond.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。