- 开始
- 2026年2月17日可信度 90%来自来源
Claude Sonnet 4.6 发布
由原文自动翻译
原标题: Claude Sonnet 4.6 Released
- “Claude Sonnet 4.6 是我们迄今最强大的 Sonnet 模型”,在编程、计算机操作、长上下文推理、智能体规划、知识型工作和设计方面全面升级[1][4]
- 该模型成为 claude.ai 和 Claude Cowork 中免费版与 Pro 用户的默认模型,免费版还新增了文件创建、连接器、技能和压缩(compaction)功能;它已登陆所有 Claude 套餐、Claude Code、API 以及所有主要云平台[1][3][4]
- 在 Claude Code 中,早期测试者约 70% 的时间更偏好它而非 Sonnet 4.5,甚至有 59% 的时间偏好它而非 Anthropic 11 月的旗舰模型 Opus 4.5,称其更不容易“过度设计”和“偷懒”[1]
- 安全研究人员发现它“性格总体温暖、诚实、有益于社会,有时还很有趣”,没有出现重大高风险失调的迹象;它比 Sonnet 4.5 更能抵御提示注入攻击,并按照 ASL-3 标准发布[1][5]
- CNBC 将此次发布与软件股抛售联系起来,IGV 软件 ETF 当年累计下跌超过 20%[3]
值得关注的功能
- 价格与 Sonnet 4.5 保持不变,为每百万输入 token 3 美元、每百万输出 token 15 美元;发布时以测试版形式提供 100 万 token 的上下文窗口,最多 128,000 个输出 token[1][2]
- Anthropic 的表格显示:SWE-bench Verified 得分 79.6%,OSWorld-Verified 得分 72.5%(Sonnet 4.5 为 61.4%),Terminal-Bench 2.0 得分 59.1%,BrowseComp 得分 74.7%[1]
- 其 OSWorld 成绩图追踪了 Sonnet 系列的发展轨迹:从 2024 年 10 月的 Claude 3.5 Sonnet 的 14.9%,一路升至 Sonnet 4.6 的 72.5%[1]
- 自适应与扩展思考,以及测试版的上下文压缩功能,可在对话接近上限时对较早的上下文进行摘要[1][2]
- 在 OfficeQA(Databricks 用于测试阅读企业图表、PDF 和表格的能力的基准)上与 Opus 4.6 表现相当;在 Vending-Bench Arena 中,它在模拟的十个月里大举投入,随后才转向追求盈利[1]
基准测试[1]
| 基准测试 | Sonnet 4.6 | Sonnet 4.5 | Opus 4.6 | Opus 4.5 | Gemini 3 Pro | GPT-5.2(所有型号) |
|---|---|---|---|---|---|---|
| 智能体终端编程(Terminal-Bench 2.0) | 59.1% | 51.0% | 65.4% | 59.8% | 56.2%(自报 54.2%) | 64.7%(Codex CLI 自报 64.0%) |
| 智能体编程(SWE-bench Verified) | 79.6% | 77.2% | 80.8% | 80.9% | 78.0%(Flash) | 80.0% |
| 智能体计算机操作(OSWorld-Verified) | 72.5% | 61.4% | 72.7% | 66.3% | — | 38.2% |
| 智能体工具使用(τ2-bench),零售 | 91.7% | 86.2% | 91.9% | 88.9% | 85.3% | 82.0% |
| 智能体工具使用(τ2-bench),电信 | 97.9% | 98.0% | 99.3% | 98.2% | 98.0% | 98.7% |
| 大规模工具使用(MCP-Atlas) | 61.3% | 43.8% | 59.5% | 62.3% | 54.1% | 60.6% |
| 智能体搜索(BrowseComp) | 74.7% | 43.9% | 84.0% | 67.8% | 59.2%(Deep Research) | 77.9%(Pro) |
| 跨学科推理(Humanity's Last Exam),不使用工具 | 33.2% | 17.7% | 40.0% | 30.8% | 37.5% | 36.6%(Pro) |
| 跨学科推理(Humanity's Last Exam),使用工具 | 49.0% | 33.6% | 53.0% | 43.4% | 45.8% | 50.0%(Pro) |
| 智能体金融分析(Finance Agent v1.1) | 63.3% | 54.5% | 60.1% | 58.8% | 55.2% | 59.0% |
| 办公任务(GDPval-AA Elo) | 1633 | 1276 | 1606 | 1416 | 1201 | 1462 |
| 新颖问题求解(ARC-AGI-2) | 58.3% | 13.6% | 68.8% | 37.6% | 31.1% | 54.2%(Pro) |
| 研究生水平推理(GPQA Diamond) | 89.9% | 83.4% | 91.3% | 87.0% | 91.9% | 93.2%(Pro) |
| 视觉推理(MMMU-Pro),不使用工具 | 74.5% | 63.4% | 73.9% | 70.6% | 81.0% | 79.5% |
| 视觉推理(MMMU-Pro),使用工具 | 75.6% | 68.9% | 77.3% | 73.9% | — | 80.4% |
| 多语言问答(MMMLU) | 89.3% | 89.5% | 91.1% | 90.8% | 91.8% | 89.6% |
参考资料 5可信度 88%总体可信度: 88%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 5 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%anthropic.com/news/claude-sonnet-4-6anthropic.com· 发布于 2026年9月28日· 开始 2026年2月17日 ✓· 占评分 31%
该文章标注日期为“2026 年 2 月 17 日”,并写道:“Claude[2] Sonnet 4.6 现已在所有 Claude 套餐、Claude Cowork、Claude Code、我们的 API 以及所有主要云平台上可用”;文档中的模型页面标注“发布日期:2026 年 2 月 17 日”,CNBC[3] 也在当天进行了报道。
- [2]90%Claude Sonnet 4.6 - Claude Platform Docsplatform.claude.com· 添加于 2026年9月28日· 占评分 31%
Anthropic's[1][5] model page: claude-sonnet-4-6, 1M-token context, 128K max output (300K Batch beta), $3/$15 per million tokens, adaptive thinking (extended deprecated), text and images in, an August 2025 reliable knowledge cutoff, "Released February 17, 2026".
- [3]85%Anthropic releases Claude Sonnet 4.6, continuing breakneck pace of AI model releasescnbc.com· 发表于 2026年2月17日· 开始 2026年2月17日· 占评分 13%
CNBC, 2026-02-17: the default for free and Pro users in Claude[2] and Claude Cowork, better at computers, coding, design and knowledge work; Anthropic's[1][5] advances had fed a sell-off in software stocks, with the iShares Expanded Tech-Software Sector ETF (IGV) down more than 20% for the year.
- [4]80%Claude Sonnet 4.6 Brings Improved Coding, Computer Use, and Office Tasksmacrumors.com· 发表于 2026年2月17日· 开始 2026年2月17日· 占评分 13%
MacRumors, 2026-02-17: "Anthropic[1][5] today updated its Sonnet model to version 4.6", available on all Claude[2] plans, with file creation, connectors, skills and compaction added for free users; Opus 4.6 stays the better choice for the hardest agentic work.
- [5]90%System Card: Claude Sonnet 4.6anthropic.com· 发表于 2026年2月17日· 占评分 13%
Anthropic's[1] system card dated "February 17, 2026", which records the model's release under the AI Safety Level 3 (ASL-3) Standard; a 6 March 2026 changelog entry updated its BrowseComp scores after an improved cheating-detection pipeline.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。