- Start
- Sep 29, 202590% CONFIDENCEfrom the source
Claude Sonnet 4.5 Released
- Anthropic called it "the best coding model in the world. It's the strongest model for building complex agents. It's the best model at using computers"[1][4]
- "Available everywhere today" as claude-sonnet-4-5 on the Claude API and in the Claude apps; Mike Krieger said it would be the default and recommended it for "basically every use case"[1][5]
- It shipped with Claude Code checkpoints and a native VS Code extension, a context editing feature and memory tool on the API, code execution and file creation in the Claude apps, and the Claude Agent SDK, the infrastructure behind Claude Code[1][4]
- "This is the most aligned frontier model we've ever released", with lower rates of sycophancy and deception, released under AI Safety Level 3 (ASL-3) protections[1][3][4]
- It arrived less than two months after Claude Opus 4.1, with Anthropic valued at $183 billion[4][5]
Notable features
- $3 per million input tokens and $15 per million output tokens, as for Sonnet 4; 200K-token context window, 64K max output and a January 2025 reliable knowledge cutoff[1][2]
- 77.2% on SWE-bench Verified, averaged over 10 trials with no test-time compute, and 82.0% with parallel test-time compute[1]
- It leads OSWorld computer use at 61.4%, where Sonnet 4 had led at 42.2% four months earlier[1]
- Anthropic observed it holding focus for more than 30 hours on complex multi-step tasks; CNBC notes Opus 4 managed seven[1][5]
- Experts in finance, law, medicine and STEM found its domain knowledge and reasoning well ahead of older models, including Opus 4.1[1]
Benchmarks[1]
| Benchmark | Claude Sonnet 4.5 | Claude Opus 4.1 | Claude Sonnet 4 | GPT-5 | Gemini 2.5 Pro |
|---|---|---|---|---|---|
| Agentic coding (SWE-bench Verified) | 77.2% / 82.0% with parallel test-time compute | 74.5% / 79.4% with parallel test-time compute | 72.7% / 80.2% with parallel test-time compute | 72.8% (GPT-5) / 74.5% (GPT-5-Codex) | 67.2% |
| Agentic terminal coding (Terminal-Bench) | 50.0% | 46.5% | 36.4% | 43.8% | 25.3% |
| Agentic tool use (τ2-bench), retail | 86.2% | 86.8% | 83.8% | 81.1% | — |
| Agentic tool use (τ2-bench), airline | 70.0% | 63.0% | 63.0% | 62.6% | — |
| Agentic tool use (τ2-bench), telecom | 98.0% | 71.5% | 49.6% | 96.7% | — |
| Computer use (OSWorld) | 61.4% | 44.4% | 42.2% | — | — |
| High school math competition (AIME 2025) | 100% (python) / 87.0% (no tools) | 78.0% | 70.5% | 99.6% (python) / 94.6% (no tools) | 88.0% |
| Graduate-level reasoning (GPQA Diamond) | 83.4% | 81.0% | 76.1% | 85.7% | 86.4% |
| Multilingual Q&A (MMMLU) | 89.1% | 89.5% | 86.5% | 89.4% | — |
| Visual reasoning (MMMU validation) | 77.8% | 77.1% | 74.4% | 84.2% | 82.0% |
| Financial analysis (Finance Agent) | 55.3% | 50.9% | 44.5% | 46.9% | 29.4% |
References 589% CONFIDENCEOverall confidence: 89%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-sonnet-4-5anthropic.com· Posted Sep 28, 2026· Starts Sep 29, 2025 ✓· 29% of score
The post is dated "Sep 29, 2025" and says "Claude[2] Sonnet 4.5 is available everywhere today"; the docs model page lists "Released September 29, 2025" and TechCrunch[4] reports the launch "On Monday".
- [2]90%Claude Sonnet 4.5 - Claude Platform Docsplatform.claude.com· Added Sep 28, 2026· 29% of score
Anthropic's[1][3] model page: claude-sonnet-4-5-20250929, 200K context window, 64K max output, $3/$15 per million tokens, extended thinking, text and images in, a January 2025 reliable knowledge cutoff, "Released September 29, 2025".
- [3]90%System Card: Claude Sonnet 4.5anthropic.com· Added Sep 28, 2026· 29% of score
Anthropic's[1] system card for Claude[2] Sonnet 4.5, "a new hybrid reasoning large language model" (September 2025), with the alignment and safety evaluations behind its ASL-3 release.
- [4]85%Anthropic launches Claude Sonnet 4.5, its best AI model for codingtechcrunch.com· Published Sep 29, 2025· Starts Sep 29, 2025· 7% of score
TechCrunch, 2025-09-29: "On Monday, Anthropic[1][3] launched a new frontier model called Claude[2] Sonnet 4.5" at Sonnet 4's $3/$15; researcher David Hershey saw it code autonomously for up to 30 hours; it arrives less than two months after Claude Opus 4.1.
- [5]85%Anthropic launches Claude Sonnet 4.5, its latest AI model that's 'more of a colleague'cnbc.com· Published Sep 29, 2025· Starts Sep 29, 2025· 7% of score
CNBC, 2025-09-29: available to all users from the $183 billion startup; Mike Krieger says it will be the default and recommends it for "basically every use case"; it runs autonomously for 30 hours against Opus 4's seven.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.