- Start
- Feb 24, 202590% CONFIDENCEfrom the source
Claude 3.7 Sonnet Released
- Anthropic called Claude 3.7 Sonnet "our most intelligent model to date and the first hybrid reasoning model on the market": one model that gives near-instant answers or "extended, step-by-step thinking that is made visible to the user"[1][4]
- It went out to every Claude plan - Free, Pro, Team and Enterprise - and to the Claude Developer Platform, Amazon Bedrock and Google Cloud's Vertex AI; extended thinking is on every surface except the free tier[1][4]
- Claude Code, Anthropic's first agentic coding tool, launched with it as a limited research preview that works from the terminal, and the GitHub integration came to every Claude plan[1]
- It cuts unnecessary refusals by 45% against its predecessor and was released under the ASL-2 standard[1][2]
- The version jumped from 3.5 to 3.7, and in February 2025 Claude 3.7 Sonnet began playing Pokémon Red live on Twitch[3][4]
Notable features
- $3 per million input tokens and $15 per million output tokens in both modes, thinking tokens included[1][4]
- API users set a thinking budget of any size up to the 128K-token output limit, trading speed and cost for quality[1]
- Anthropic's charts give 62.3% on SWE-bench Verified (70.3% with a custom scaffold) against 49.0% for the upgraded 3.5 Sonnet, and 81.2% (retail) and 58.4% (airline) on TAU-bench[1]
- Anthropic says it optimized less for maths and coding competitions and more for real-world business tasks[1]
- In standard mode it is an upgraded Claude 3.5 Sonnet; in extended thinking mode it reflects before answering, which helps maths, physics, instruction following and coding[1]
Benchmarks[1]
| Benchmark | Claude 3.7 Sonnet (64K extended thinking) | Claude 3.7 Sonnet (no extended thinking) | Claude 3.5 Sonnet (new) | OpenAI o1¹ | OpenAI o3-mini¹ (high) | DeepSeek R1 (32K extended thinking) | Grok 3 Beta (extended thinking) |
|---|---|---|---|---|---|---|---|
| Graduate-level reasoning (GPQA Diamond)³ | 78.2% / 84.8% | 68.0% | 65.0% | 75.7% / 78.0% | 79.7% | 71.5% | 80.2% / 84.6% |
| Agentic coding (SWE-bench Verified)² | — | 62.3% / 70.3% | 49.0% | 48.9% | 49.3% | 49.2% | — |
| Agentic tool use (TAU-bench), retail | — | 81.2% | 71.5% | 73.5% | — | — | — |
| Agentic tool use (TAU-bench), airline | — | 58.4% | 48.8% | 54.2% | — | — | — |
| Multilingual Q&A (MMMLU) | 86.1% | 83.2% | 82.1% | 87.7% | 79.5% | — | — |
| Visual reasoning (MMMU validation) | 75% | 71.8% | 70.4% | 78.2% | — | — | 76.0% / 78.0% |
| Instruction-following (IFEval) | 93.2% | 90.8% | 90.2% | — | — | 83.3% | — |
| Math problem-solving (MATH 500) | 96.2% | 82.2% | 78.0% | 96.4% | 97.9% | 97.3% | — |
| High school math competition (AIME 2024)³ | 61.3% / 80.0% | 23.3% | 16.0% | 79.2% / 83.3% | 87.3% | 79.8% | 83.9% / 93.3% |
Pass@1 averaged over several trials (up to 16 for AIME and SWE-bench Verified); a second figure benefits from parallel test-time compute. ¹ o1 TAU-bench results were reported by OpenAI and later deleted; TAU-bench results may not be comparable. ² 62.3% is pass@1 on 500 problems with bash/editor tools plus a "thinking tool"; 70.3% uses internal scoring and a custom scaffold on a reduced subset; DeepSeek R1 uses the Agentless framework. ³ Claude 3.7 Sonnet's high GPQA and AIME 2024 scores use internal scoring with parallel test-time compute; o1's and Grok 3's use majority voting with 64 samples.
References 485% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 4 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-3-7-sonnetanthropic.com· Posted Sep 28, 2026· Starts Feb 24, 2025 ✓· 32% of score
The post is dated "Feb 24, 2025": "Today, we're announcing Claude 3.7 Sonnet" and it "is now available on all Claude plans"; TechCrunch[4] says it "is rolling out to all users and developers on Monday" (24 February).
- [2]90%Claude 3.7 Sonnet System Cardanthropic.com· Added Sep 28, 2026· 32% of score
Anthropic's[1] system card for "a hybrid reasoning model": it explains why users see the model's thinking, and states "Claude 3.7 Sonnet is released under the ASL-2 standard".
- [3]75%Claude (AI) - Wikipediaen.wikipedia.org· Added Sep 28, 2026· 32% of score
The Sonnet table dates Claude 3.7 Sonnet to 24 February 2025; the article adds that in February 2025 Claude 3.7 Sonnet playing Pokémon Red began streaming on Twitch to thousands of viewers.[1]
- [4]85%Anthropic launches a new AI model that 'thinks' as long as you wanttechcrunch.com· Published Feb 24, 2025· Starts Feb 24, 2025· 3% of score
TechCrunch, 2025-02-24: the model "is rolling out to all users and developers on Monday"; free users get the standard mode, only paid plans the reasoning; $3/$15 per million tokens against o3-mini's $1.10/$4.40 and DeepSeek R1's $0.55/$2.19; "(Yes, the company skipped a number.)"
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.