- 开始
- 2024年10月22日可信度 90%来自来源
升级版 Claude 3.5 Sonnet 发布
由原文自动翻译
原标题: Claude 3.5 Sonnet Upgrade Released
- Anthropic 宣布“升级版 Claude 3.5 Sonnet”,称其“相较前代在各方面均有提升,编程能力的进步尤为显著”,当天即面向所有用户开放,并预告新款 Claude 3.5 Haiku 将于本月晚些时候推出[1]
- 计算机操作功能进入公开测试版:Claude 会查看屏幕、移动光标、点击按钮并输入文字,Anthropic 称 3.5 Sonnet 是“首个在公开测试版中提供计算机操作功能的前沿人工智能模型”——“仍处于实验阶段——有时会显得笨拙且容易出错”[1][4]
- 为了操控光标,Claude 会在每张截图上计算需要移动的像素数;Anthropic 用计算器、文本编辑器等简单软件来训练这项技能[3]
- 开发者当天即可在 Anthropic API、Amazon Bedrock 和 Google Cloud 的 Vertex AI 上使用该功能;Asana、Canva、Cognition、DoorDash、Replit 和 The Browser Company 是早期用户,Amazon 也获得了提前使用权[1][5]
- 美国和英国的人工智能安全研究所在发布前联合对其进行了测试;Anthropic 将其保持在 ASL-2 级别,并新增了可识别计算机操作以及是否正在发生危害的分类器[1]
值得关注的功能
- 价格和速度与 6 月版 3.5 Sonnet 相同[1]
- SWE-bench Verified 从 33.4% 提升到 49.0%,超过了包括 OpenAI o1-preview 之外的所有公开可用模型[1]
- TAU-bench 智能体工具使用得分在零售领域从 62.6% 提升到 69.2%,在航空领域从 36.0% 提升到 46.0%[1]
- 在 OSWorld 上,仅凭截图得分 14.9%,高于次优系统的 7.8%;在允许更多步骤的情况下得分 22.0%[1][2]
- 滚动、拖拽和缩放操作对它来说仍然很难,Anthropic 建议先从低风险任务开始尝试[1]
基准测试[1]
| 基准测试 | Claude 3.5 Sonnet(新版) | Claude 3.5 Haiku | Claude 3.5 Sonnet | GPT-4o* | GPT-4o mini* | Gemini 1.5 Pro | Gemini 1.5 Flash |
|---|---|---|---|---|---|---|---|
| 研究生水平推理(GPQA Diamond) | 65.0%(0-shot CoT) | 41.6%(0-shot CoT) | 59.4%(0-shot CoT) | 53.6%(0-shot CoT) | 40.2%(0-shot CoT) | 59.1%(0-shot CoT) | 51.0%(0-shot CoT) |
| 本科水平知识(MMLU Pro) | 78.0%(0-shot CoT) | 65.0%(0-shot CoT) | 75.1%(0-shot CoT) | — | — | 75.8%(0-shot CoT) | 67.3%(0-shot CoT) |
| 代码(HumanEval) | 93.7%(0-shot) | 88.1%(0-shot) | 92.0%(0-shot) | 90.2%(0-shot) | 87.2%(0-shot) | — | — |
| 数学问题求解(MATH) | 78.3%(0-shot CoT) | 69.2%(0-shot CoT) | 71.1%(0-shot CoT) | 76.6%(0-shot CoT) | 70.2%(0-shot CoT) | 86.5%(4-shot CoT) | 77.9%(4-shot CoT) |
| 高中数学竞赛(AIME 2024) | 16.0%(0-shot CoT) | 5.3%(0-shot CoT) | 9.6%(0-shot CoT) | 9.3%(0-shot CoT) | — | — | — |
| 视觉问答(MMMU) | 70.4%(0-shot CoT) | — | 68.3%(0-shot CoT) | 69.1%(0-shot CoT) | 59.4%(0-shot CoT) | 65.9%(0-shot CoT) | 62.3%(0-shot CoT) |
| 智能体编程(SWE-bench Verified) | 49.0% | 40.6% | 33.4% | — | — | — | — |
| 智能体工具使用(TAU-bench),零售 | 69.2% | 51.0% | 62.6% | — | — | — | — |
| 智能体工具使用(TAU-bench),航空 | 46.0% | 22.8% | 36.0% | — | — | — | — |
* Anthropic 的评测表未纳入 OpenAI 的 o1 模型系列,因为该系列依赖大量的响应前计算时间,不同于典型模型。
参考资料 5可信度 90%总体可信度: 90%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 5 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/news/3-5-models-and-computer-useanthropic.com· 发布于 2026年9月28日· 开始 2024年10月22日 ✓· 占评分 46%
该文章标注日期为“2024 年 10 月 22 日”,并写道:“升级版 Claude 3.5 Sonnet 现已面向所有用户开放。从今天起,开发者可以使用计算机操作测试版进行构建”;TechCrunch[4] 和 CNBC[5] 均在同一天报道该发布“于周二”进行。
- [2]90%Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnetassets.anthropic.com· 添加于 2026年9月28日· 占评分 46%
Anthropic's[1][3] model card addendum: the upgraded 3.5 Sonnet sets new state-of-the-art results in agentic coding (SWE-bench Verified), agentic tasks (TAU-bench) and computer use from screenshots (OSWorld), with a 14.9% average OSWorld success rate from screenshots alone.
- [3]90%Developing a computer use modelanthropic.com· 发表于 2024年10月22日· 占评分 3%
Anthropic's[1][2] research post on the skill: Claude looks at screenshots and "counts how many pixels vertically or horizontally it needs to move a cursor in order to click in the correct place", and was trained on simple software such as a calculator and a text editor without internet access.
- [4]85%Anthropic's new AI model can control your PCtechcrunch.com· 发表于 2024年10月22日· 开始 2024年10月22日· 占评分 3%
TechCrunch, 2024-10-22: "Anthropic[1][2][3] on Tuesday released an upgraded version of its Claude 3.5 Sonnet model that can understand and interact with any desktop app" through a Computer Use API in open beta on the API, Bedrock and Vertex AI.
- [5]85%Amazon-backed Anthropic debuts AI agents that can do complex tasks, racing against OpenAI, Microsoft and Googlecnbc.com· 发表于 2024年10月22日· 开始 2024年10月22日· 占评分 3%
CNBC, 2024-10-22: chief science officer Jared Kaplan says it can do tasks with "tens or even hundreds of steps"; Amazon had early access and beta testers included Asana, Canva and Notion; released Tuesday in public beta for developers.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。