- 开始
- 2024年6月20日可信度 92%来自techcrunch.com
Claude 3.5 Sonnet 发布
由原文自动翻译
原标题: Claude 3.5 Sonnet Released
- “今天,我们发布 Claude 3.5 Sonnet——我们即将推出的 Claude 3.5 模型系列的首个版本”,其表现超越了自家的 Claude 3 Opus,“同时保持中端模型 Claude 3 Sonnet 的速度和成本”[1][4]
- 在 Claude.ai 和 Claude iOS 应用上免费提供,Pro 和 Team 订阅用户享有更高的速率限制,并同日在 Anthropic API、Amazon Bedrock 和 Google Cloud 的 Vertex AI 上线[1][5]
- 该模型随附 Artifacts 功能,在对话旁边打开一个窗口,展示代码、文档和网站设计,方便用户实时查看、编辑并在其基础上继续构建[1][3]
- 该模型仍处于 ASL-2 级别;Anthropic 将其提供给英国人工智能安全研究所进行部署前测试,该研究所又将结果分享给了美国人工智能安全研究所[1]
- Anthropic 表示,将在 2024 年晚些时候补齐 Claude 3.5 Haiku 和 Claude 3.5 Opus,完善整个系列[1]
值得关注的功能
- 定价为每百万输入 token 3 美元、每百万输出 token 15 美元,上下文窗口为 20 万 token;AWS 称其比 Claude 3 Opus 便宜 80%[1][5]
- 运行速度是 Claude 3 Opus 的两倍[1][4]
- Anthropic 的模型卡:GPQA Diamond 得分 59.4%,MMLU 得分 88.7%,HumanEval 得分 92.0%,MATH 得分 71.1%,相较 Claude 3 Opus 的 50.4%、86.8%、84.9% 和 60.1%[2]
- 在一项内部智能体编程评测(为开源代码库修复漏洞或添加功能)中解决了 64% 的问题,而 Claude 3 Opus 为 38%[1][2]
- 迄今最强的视觉模型:更擅长阅读图表和图形,也更擅长转录不清晰图像中的文字,MMMU 得分 68.3%[1][2]
基准测试[1]
| 基准测试 | Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b(早期快照) |
|---|---|---|---|---|---|
| 研究生水平推理(GPQA, Diamond) | 59.4%*(0-shot CoT) | 50.4%(0-shot CoT) | 53.6%(0-shot CoT) | — | — |
| 本科水平知识(MMLU),5-shot | 88.7%**(5-shot) | 86.8%(5-shot) | — | 85.9%(5-shot) | 86.1%(5-shot) |
| 本科水平知识(MMLU),0-shot CoT | 88.3%(0-shot CoT) | 85.7%(0-shot CoT) | 88.7%(0-shot CoT) | — | — |
| 代码(HumanEval) | 92.0%(0-shot) | 84.9%(0-shot) | 90.2%(0-shot) | 84.1%(0-shot) | 84.1%(0-shot) |
| 多语言数学(MGSM) | 91.6%(0-shot CoT) | 90.7%(0-shot CoT) | 90.5%(0-shot CoT) | 87.5%(8-shot) | — |
| 文本推理(DROP, F1 score) | 87.1(3-shot) | 83.1(3-shot) | 83.4(3-shot) | 74.9(可变样本数) | 83.5(3-shot,预训练模型) |
| 综合评测(BIG-Bench-Hard) | 93.1%(3-shot CoT) | 86.8%(3-shot CoT) | — | 89.2%(3-shot CoT) | 85.3%(3-shot CoT,预训练模型) |
| 数学问题求解(MATH) | 71.1%(0-shot CoT) | 60.1%(0-shot CoT) | 76.6%(0-shot CoT) | 67.7%(4-shot) | 57.8%(4-shot CoT) |
| 小学数学(GSM8K) | 96.4%(0-shot CoT) | 95.0%(0-shot CoT) | — | 90.8%(11-shot) | 94.1%(8-shot CoT) |
* Claude 3.5 Sonnet 在 5-shot CoT、maj@32 条件下的 GPQA 得分为 67.2%。** Claude 3.5 Sonnet 在 5-shot CoT 提示下的 MMLU 得分为 90.4%。
参考资料 5可信度 85%总体可信度: 85%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 5 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/news/claude-3-5-sonnetanthropic.com· 发布于 2026年9月28日· 开始 2024年6月20日· 占评分 32%
Anthropic[2] 写道:“今天,我们发布 Claude 3.5 Sonnet”,并且它“现已在 Claude.ai 上免费提供”;该页面的标注日期现显示为 2024 年 6 月 21 日,但 TechCrunch[4] 和 AWS 均将发布日期报道为 2024 年 6 月 20 日,且模型 ID 为 claude-3-5-sonnet-20240620。
- [2]90%Claude 3.5 Sonnet Model Card Addendumwww-cdn.anthropic.com· 添加于 2026年9月28日· 占评分 32%
Anthropic's[1] model card addendum: 59.4% on GPQA Diamond, 88.7% on MMLU (5-shot), 92.0% on HumanEval, 71.1% on MATH and 68.3% on MMMU, against Claude 3 Opus's 50.4%, 86.8%, 84.9%, 60.1% and 59.4%; 64% on the internal agentic coding evaluation against Opus's 38%.
- [3]75%Claude (AI) - Wikipediaen.wikipedia.org· 添加于 2026年9月28日· 占评分 32%
"On June 20, 2024, Anthropic[1][2] released Claude 3.5 Sonnet, which, according to the company's own benchmarks, performed better than the larger Claude 3 Opus", alongside Artifacts, which previews code, SVG graphics or websites in a separate window.
- [4]92%Anthropic claims its latest model is best-in-classtechcrunch.com· 发表于 2024年6月20日· 开始 2024年6月20日 ✓· 占评分 1%
TechCrunch, published 2024-06-20T14:00Z: Anthropic[1][2] "is releasing a powerful new generative AI model called Claude 3.5 Sonnet", about twice the speed of Claude 3 Opus, alongside Artifacts. It dates the launch 20 June, where Anthropic's own header now shows the 21 June UTC timestamp.
- [5]90%Anthropic's Claude 3.5 Sonnet model now available in Amazon Bedrock: Even more intelligence than Claude 3 Opus at one-fifth the costaws.amazon.com· 发表于 2024年6月20日· 开始 2024年6月20日· 占评分 1%
AWS News Blog, 20 JUN 2024: "Today, Anthropic[1][2] introduced Claude 3.5 Sonnet ... now available in Amazon Bedrock"; 80 percent cheaper than Opus and about 10 percent better than Claude 3 Opus on most vision benchmarks.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。