- 开始
- 2024年3月4日可信度 95%来自anthropic.com
Claude 3 Sonnet 发布
由原文自动翻译
原标题: Claude 3 Sonnet Released
- Anthropic 于 2024 年 3 月 4 日宣布推出 Claude 3 系列——Haiku、Sonnet 和 Opus,Opus 和 Sonnet“现已在 claude.ai 和 Claude API 中可用”,该 API 已全面开放至 159 个国家;Haiku 稍后跟进[3]
- Sonnet 为“claude.ai 的免费体验”提供支持,Opus 则保留给 Claude Pro 订阅用户,并同日登陆 Amazon Bedrock,以及在 Google Cloud 的 Vertex AI Model Garden 上进入私有预览[3][4]
- Anthropic 将其定位为“在智能和速度之间取得理想平衡——尤其适合企业工作负载”的模型,在大多数工作负载上速度是 Claude 2 和 Claude 2.1 的两倍[3][4]
- 该系列在 Anthropic 的《负责任扩展政策》下维持在 AI 安全等级 2(ASL-2),并且相比早期 Claude 模型,在接近其防护边界时拒绝无害提问的情况要少得多[3]
- Anthropic 于 2025 年 7 月退役 Claude 3 Sonnet 时,约有 200 人在旧金山聚集参加了一场“葬礼”[2]
值得关注的功能
- 定价为每百万输入 token 3 美元、每百万输出 token 15 美元,上下文窗口为 20 万 token[3]
- 首个具备视觉能力的 Claude 世代:可读取照片、图表、图形和技术示意图,单次请求最多 20 张图片,不过 Anthropic 阻止了它识别人物身份[1][5]
- 模型卡中的基准测试成绩:MMLU(5-shot)得分 79.0%,GPQA Diamond 得分 40.4%,GSM8K 得分 92.3%,HumanEval Python 编程得分 73.0%[1]
- 使用 PyTorch、JAX 和 Triton,在 Amazon Web Services 和 Google Cloud 的硬件上训练;其回答基于 2023 年 8 月之前的数据,且无法搜索网络[1][5]
- Anthropic 建议的用途:大型知识库检索、产品推荐与预测、代码生成以及从图像中解析文本[3]
基准测试[3]
| 基准测试 | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| 本科水平知识(MMLU) | 86.8%(5-shot) | 79.0%(5-shot) | 75.2%(5-shot) | 86.4%(5-shot) | 70.0%(5-shot) | 83.7%(5-shot) | 71.8%(5-shot) |
| 研究生水平推理(GPQA, Diamond) | 50.4%(0-shot CoT) | 40.4%(0-shot CoT) | 33.3%(0-shot CoT) | 35.7%(0-shot CoT) | 28.1%(0-shot CoT) | — | — |
| 小学数学(GSM8K) | 95.0%(0-shot CoT) | 92.3%(0-shot CoT) | 88.9%(0-shot CoT) | 92.0%(5-shot CoT) | 57.1%(5-shot) | 94.4%(Maj1@32) | 86.5%(Maj1@32) |
| 数学问题求解(MATH) | 60.1%(0-shot CoT) | 43.1%(0-shot CoT) | 38.9%(0-shot CoT) | 52.9%(4-shot) | 34.1%(4-shot) | 53.2%(4-shot) | 32.6%(4-shot) |
| 多语言数学(MGSM) | 90.7%(0-shot) | 83.5%(0-shot) | 75.1%(0-shot) | 74.5%(8-shot) | — | 79.0%(8-shot) | 63.5%(8-shot) |
| 代码(HumanEval) | 84.9%(0-shot) | 73.0%(0-shot) | 75.9%(0-shot) | 67.0%(0-shot) | 48.1%(0-shot) | 74.4%(0-shot) | 67.7%(0-shot) |
| 文本推理(DROP, F1 score) | 83.1(3-shot) | 78.9(3-shot) | 78.4(3-shot) | 80.9(3-shot) | 64.1(3-shot) | 82.4(可变样本数) | 74.1(可变样本数) |
| 综合评测(BIG-Bench-Hard) | 86.8%(3-shot CoT) | 82.9%(3-shot CoT) | 73.7%(3-shot CoT) | 83.1%(3-shot CoT) | 66.6%(3-shot CoT) | 83.6%(3-shot CoT) | 75.0%(3-shot CoT) |
| 知识问答(ARC-Challenge) | 96.4%(25-shot) | 93.2%(25-shot) | 89.2%(25-shot) | 96.3%(25-shot) | 85.2%(25-shot) | — | — |
| 常识(HellaSwag) | 95.4%(10-shot) | 89.0%(10-shot) | 85.9%(10-shot) | 95.3%(10-shot) | 85.5%(10-shot) | 87.8%(10-shot) | 84.7%(10-shot) |
参考资料 5可信度 83%总体可信度: 83%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 5 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/claude-3-model-cardanthropic.com· 发布于 2026年9月28日· 开始 2024年3月4日· 占评分 48%
模型卡中未标注发布日期,因此以发布文章的日期为准:“新一代 Claude 隆重登场,2024 年 3 月 4 日”——“Opus 和 Sonnet 现已在 claude.ai 和 Claude API 中可用”;AWS 于 2024 年 3 月 4 日发布的文章宣布“Anthropic[3] 的 Claude 3 Sonnet 今日在 Amazon[4] Bedrock 上可用”。
- [2]75%Claude (AI) - Wikipediaen.wikipedia.org· 添加于 2026年9月28日· 占评分 48%
The Sonnet table lists Claude 3 Sonnet as released 4 March 2024 and discontinued; the article adds that when Anthropic[1][3] retired it in July 2025 around 200 people gathered in San Francisco for a "funeral".
- [3]95%Introducing the next generation of Claudeanthropic.com· 发表于 2024年3月4日· 开始 2024年3月4日 ✓· 占评分 1%
Anthropic's[1] launch post, dated "Mar 4, 2024": "Opus and Sonnet are now available to use in claude.ai and the Claude API which is now generally available in 159 countries", with Sonnet at $3/$15 per million tokens and a 200K context window, powering the free claude.ai. The model card that is the source carries no launch date, so this row dates the pin.
- [4]90%Anthropic's Claude 3 Sonnet foundation model is now available in Amazon Bedrockaws.amazon.com· 发表于 2024年3月4日· 开始 2024年3月4日· 占评分 1%
AWS News Blog post by Channy Yun, 04 MAR 2024: "We're also announcing the availability of Anthropic's[1][3] Claude 3 Sonnet today in Amazon Bedrock, with Claude 3 Opus and Claude 3 Haiku coming soon"; Sonnet is two times faster than Claude 2 and 2.1.
- [5]85%Anthropic claims its new AI chatbot models beat OpenAI's GPT-4techcrunch.com· 发表于 2024年3月4日· 开始 2024年3月4日· 占评分 1%
TechCrunch's same-day report: Opus and Sonnet "are available now on the web and via Anthropic's[1][3] dev console and API, Amazon's[4] Bedrock platform and Google's Vertex AI"; Claude 3 is Anthropic's first multimodal model, takes up to 20 images per request, cannot identify people and answers only from data before August 2023.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。