- 开始
- 2024年3月4日可信度 90%来自来源
Claude 3 Opus 发布
由原文自动翻译
原标题: Claude 3 Opus Released
- Anthropic 于 2024 年 3 月 4 日宣布推出 Claude 3 系列——Haiku、Sonnet 和 Opus,“按能力从低到高排列”——Opus 和 Sonnet 当天即在 claude.ai 和 Claude API 上可用,现已在 159 个国家全面开放[1][4]
- 在 claude.ai 上,Opus 面向 Claude Pro 订阅用户,而 Sonnet 为免费版提供支持;Sonnet 当天登陆 Amazon Bedrock,并在 Google Cloud Vertex AI 上进入私有预览,“Opus 和 Haiku 也将很快登陆这两个平台”[1]
- 在红队测试发现“目前发生灾难性风险的可能性微乎其微”后,Anthropic 将该模型保持在 AI 安全等级 2(ASL-2)[1]
- CNBC 指出,Anthropic 的投资方包括 Google、Salesforce 和 Amazon,此前一年的融资总额约为 73 亿美元[4]
- Claude 3 Opus 于 2025 年 6 月 30 日被弃用,并于 2026 年 1 月 5 日从 Claude API 中正式退役,由 Claude Opus 4.8 取代[3]
值得关注的功能
- Claude 3 Opus 是“我们最智能的模型,在高度复杂的任务上拥有市场领先的表现”,定价为每百万输入 token 15 美元、每百万输出 token 75 美元[1][5]
- Anthropic 公布的基准测试成绩:MMLU 得分 86.8%,GPQA Diamond 得分 50.4%,GSM8K 得分 95.0%,HumanEval 得分 84.9%,均领先于 GPT-4 的 86.4%、35.7%、92.0% 和 67.0%[1]
- 发布时的上下文窗口为 20 万 token,超过 100 万 token 的输入“可用于特定用例”;在“大海捞针”召回测试中,准确率超过 99%,有时还能察觉出植入的句子看起来是人为插入的[1][5]
- Claude 首款视觉模型:可读取照片、图表、图形和技术示意图,单次请求最多 20 张图片,但不会识别照片中的人物身份[1][5]
- Opus“明显更不容易”像早期 Claude 模型那样拒绝无害的提问,在困难的开放式事实性问题上的准确率是 Claude 2.1 的两倍;知识截止日期为 2023 年 8 月[1][2]
基准测试[1]
| 基准测试 | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| MMLU(本科水平知识) | 86.8% 5-shot | 79.0% 5-shot | 75.2% 5-shot | 86.4% 5-shot | 70.0% 5-shot | 83.7% 5-shot | 71.8% 5-shot |
| GPQA, Diamond(研究生水平推理) | 50.4% 0-shot CoT | 40.4% 0-shot CoT | 33.3% 0-shot CoT | 35.7% 0-shot CoT | 28.1% 0-shot CoT | — | — |
| GSM8K(小学数学) | 95.0% 0-shot CoT | 92.3% 0-shot CoT | 88.9% 0-shot CoT | 92.0% 5-shot CoT | 57.1% 5-shot | 94.4% Maj1@32 | 86.5% Maj1@32 |
| MATH(数学问题求解) | 60.1% 0-shot CoT | 43.1% 0-shot CoT | 38.9% 0-shot CoT | 52.9% 4-shot | 34.1% 4-shot | 53.2% 4-shot | 32.6% 4-shot |
| MGSM(多语言数学) | 90.7% 0-shot | 83.5% 0-shot | 75.1% 0-shot | 74.5% 8-shot | — | 79.0% 8-shot | 63.5% 8-shot |
| HumanEval(代码) | 84.9% 0-shot | 73.0% 0-shot | 75.9% 0-shot | 67.0% 0-shot | 48.1% 0-shot | 74.4% 0-shot | 67.7% 0-shot |
| DROP, F1 score(文本推理) | 83.1 3-shot | 78.9 3-shot | 78.4 3-shot | 80.9 3-shot | 64.1 3-shot | 82.4 可变样本数 | 74.1 可变样本数 |
| BIG-Bench-Hard(综合评测) | 86.8% 3-shot CoT | 82.9% 3-shot CoT | 73.7% 3-shot CoT | 83.1% 3-shot CoT | 66.6% 3-shot CoT | 83.6% 3-shot CoT | 75.0% 3-shot CoT |
| ARC-Challenge(知识问答) | 96.4% 25-shot | 93.2% 25-shot | 89.2% 25-shot | 96.3% 25-shot | 85.2% 25-shot | — | — |
| HellaSwag(常识) | 95.4% 10-shot | 89.0% 10-shot | 85.9% 10-shot | 95.3% 10-shot | 85.5% 10-shot | 87.8% 10-shot | 84.7% 10-shot |
参考资料 5可信度 90%总体可信度: 90%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 5 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/news/claude-3-familyanthropic.com· 发布于 2026年9月28日· 开始 2024年3月4日 ✓· 占评分 33%
Anthropic[2] 的发布文章标注日期为“2024 年 3 月 4 日”,并写道“Opus 和 Sonnet 现已在 claude[3].ai 和 Claude API 中可用,该 API 现已在 159 个国家全面开放”;CNBC[4] 和 TechCrunch[5] 均报道该发布发生在同一个周一。
- [2]90%The Claude 3 Model Family: Opus, Sonnet, Haiku (model card)anthropic.com· 添加于 2026年9月28日· 占评分 33%
Anthropic's[1] Claude[3] 3 model card (the link resolves to the PDF): the Claude 3 models' knowledge cutoff is August 2023 and they are offered through the Claude API, Amazon Bedrock and Google Vertex AI; it carries the full evaluations behind the launch post's table.
- [3]90%Model deprecations - Claude Platform Docsplatform.claude.com· 添加于 2026年9月28日· 占评分 33%
Anthropic's[1][2] deprecation history: on 30 June 2025 it notified developers of Claude Opus 3's retirement, and claude-3-opus-20240229 was retired on 5 January 2026 with claude-opus-4-8 as the recommended replacement.
- [4]85%Anthropic, backed by Amazon and Google, debuts its most powerful chatbot yetcnbc.com· 发表于 2024年3月4日· 开始 2024年3月4日· 占评分 1%
CNBC's same-day report (published 2024-03-04T14:00Z): "Anthropic[1][2] on Monday debuted Claude[3] 3"; Opus outperformed GPT-4 and Gemini Ultra on benchmark tests, it is Anthropic's first multimodal model, Opus and Sonnet are available in 159 countries, and Anthropic's backers include Google, Salesforce and Amazon after about $7.3 billion in funding deals over the past year.
- [5]85%Anthropic claims its new AI chatbot models beat OpenAI's GPT-4techcrunch.com· 发表于 2024年3月4日· 开始 2024年3月4日· 占评分 1%
TechCrunch (Kyle Wiggers, 11:50 AM PST, 4 March 2024): Claude[3] 3 is Anthropic's[1][2] first multimodal model, can analyze up to 20 images in one request, starts with a 200,000-token context window (1 million for select customers), answers from data before August 2023, and Opus costs $15 per million input and $75 per million output tokens.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。