- Start
- Feb 17, 202690% CONFIDENCEfrom the source
Claude Sonnet 4.6 Released
- "Claude Sonnet 4.6 is our most capable Sonnet model yet", a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work and design[1][4]
- It became the default model in claude.ai and Claude Cowork for Free and Pro users, and the free tier gained file creation, connectors, skills and compaction; it is on every Claude plan, Claude Code, the API and all major cloud platforms[1][3][4]
- In Claude Code, early testers preferred it to Sonnet 4.5 about 70% of the time and even to Opus 4.5, Anthropic's November flagship, 59% of the time, calling it less prone to overengineering and "laziness"[1]
- Safety researchers found "a broadly warm, honest, prosocial, and at times funny character" and no signs of major high-stakes misalignment; it resists prompt injection much better than Sonnet 4.5 and was released under ASL-3[1][5]
- CNBC tied the launch to a sell-off in software stocks, with the IGV software ETF down more than 20% for the year[3]
Notable features
- Price unchanged from Sonnet 4.5 at $3 per million input tokens and $15 per million output tokens; a 1M-token context window (in beta at launch) and 128K max output[1][2]
- Anthropic's table: 79.6% on SWE-bench Verified, 72.5% on OSWorld-Verified (Sonnet 4.5: 61.4%), 59.1% on Terminal-Bench 2.0 and 74.7% on BrowseComp[1]
- Its OSWorld chart traces the Sonnet line from 14.9% for the October 2024 Claude 3.5 Sonnet to 72.5% for Sonnet 4.6[1]
- Adaptive and extended thinking, plus context compaction in beta that summarises older context as a conversation nears its limit[1][2]
- It matches Opus 4.6 on OfficeQA, Databricks' test of reading enterprise charts, PDFs and tables, and in Vending-Bench Arena it invested heavily for ten simulated months before pivoting to profit[1]
Benchmarks[1]
| Benchmark | Sonnet 4.6 | Sonnet 4.5 | Opus 4.6 | Opus 4.5 | Gemini 3 Pro | GPT-5.2 (all models) |
|---|---|---|---|---|---|---|
| Agentic terminal coding (Terminal-Bench 2.0) | 59.1% | 51.0% | 65.4% | 59.8% | 56.2% (54.2% self-reported) | 64.7% (64.0% self-reported, Codex CLI) |
| Agentic coding (SWE-bench Verified) | 79.6% | 77.2% | 80.8% | 80.9% | 78.0% (Flash) | 80.0% |
| Agentic computer use (OSWorld-Verified) | 72.5% | 61.4% | 72.7% | 66.3% | — | 38.2% |
| Agentic tool use (τ2-bench), retail | 91.7% | 86.2% | 91.9% | 88.9% | 85.3% | 82.0% |
| Agentic tool use (τ2-bench), telecom | 97.9% | 98.0% | 99.3% | 98.2% | 98.0% | 98.7% |
| Scaled tool use (MCP-Atlas) | 61.3% | 43.8% | 59.5% | 62.3% | 54.1% | 60.6% |
| Agentic search (BrowseComp) | 74.7% | 43.9% | 84.0% | 67.8% | 59.2% (Deep Research) | 77.9% (Pro) |
| Multidisciplinary reasoning (Humanity's Last Exam), without tools | 33.2% | 17.7% | 40.0% | 30.8% | 37.5% | 36.6% (Pro) |
| Multidisciplinary reasoning (Humanity's Last Exam), with tools | 49.0% | 33.6% | 53.0% | 43.4% | 45.8% | 50.0% (Pro) |
| Agentic financial analysis (Finance Agent v1.1) | 63.3% | 54.5% | 60.1% | 58.8% | 55.2% | 59.0% |
| Office tasks (GDPval-AA Elo) | 1633 | 1276 | 1606 | 1416 | 1201 | 1462 |
| Novel problem-solving (ARC-AGI-2) | 58.3% | 13.6% | 68.8% | 37.6% | 31.1% | 54.2% (Pro) |
| Graduate-level reasoning (GPQA Diamond) | 89.9% | 83.4% | 91.3% | 87.0% | 91.9% | 93.2% (Pro) |
| Visual reasoning (MMMU-Pro), without tools | 74.5% | 63.4% | 73.9% | 70.6% | 81.0% | 79.5% |
| Visual reasoning (MMMU-Pro), with tools | 75.6% | 68.9% | 77.3% | 73.9% | — | 80.4% |
| Multilingual Q&A (MMMLU) | 89.3% | 89.5% | 91.1% | 90.8% | 91.8% | 89.6% |
References 588% CONFIDENCEOverall confidence: 88%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-sonnet-4-6anthropic.com· Posted Sep 28, 2026· Starts Feb 17, 2026 ✓· 31% of score
The post is dated "Feb 17, 2026" and says "Claude[2] Sonnet 4.6 is available now on all Claude plans, Claude Cowork, Claude Code, our API, and all major cloud platforms"; the docs model page lists "Released February 17, 2026" and CNBC[3] reported it that day.
- [2]90%Claude Sonnet 4.6 - Claude Platform Docsplatform.claude.com· Added Sep 28, 2026· 31% of score
Anthropic's[1][5] model page: claude-sonnet-4-6, 1M-token context, 128K max output (300K Batch beta), $3/$15 per million tokens, adaptive thinking (extended deprecated), text and images in, an August 2025 reliable knowledge cutoff, "Released February 17, 2026".
- [3]85%Anthropic releases Claude Sonnet 4.6, continuing breakneck pace of AI model releasescnbc.com· Published Feb 17, 2026· Starts Feb 17, 2026· 13% of score
CNBC, 2026-02-17: the default for free and Pro users in Claude[2] and Claude Cowork, better at computers, coding, design and knowledge work; Anthropic's[1][5] advances had fed a sell-off in software stocks, with the iShares Expanded Tech-Software Sector ETF (IGV) down more than 20% for the year.
- [4]80%Claude Sonnet 4.6 Brings Improved Coding, Computer Use, and Office Tasksmacrumors.com· Published Feb 17, 2026· Starts Feb 17, 2026· 13% of score
MacRumors, 2026-02-17: "Anthropic[1][5] today updated its Sonnet model to version 4.6", available on all Claude[2] plans, with file creation, connectors, skills and compaction added for free users; Opus 4.6 stays the better choice for the hardest agentic work.
- [5]90%System Card: Claude Sonnet 4.6anthropic.com· Published Feb 17, 2026· 13% of score
Anthropic's[1] system card dated "February 17, 2026", which records the model's release under the AI Safety Level 3 (ASL-3) Standard; a 6 March 2026 changelog entry updated its BrowseComp scores after an improved cheating-detection pipeline.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.