- Start
- Jun 20, 202492% CONFIDENCEfrom techcrunch.com
Claude 3.5 Sonnet Released
- "Today, we're launching Claude 3.5 Sonnet—our first release in the forthcoming Claude 3.5 model family", which outperforms Claude 3 Opus "with the speed and cost of our mid-tier model, Claude 3 Sonnet"[1][4]
- Free on Claude.ai and the Claude iOS app, with higher rate limits for Pro and Team subscribers, and available the same day on the Anthropic API, Amazon Bedrock and Google Cloud's Vertex AI[1][5]
- It launched with Artifacts, a window beside the conversation where code, documents and website designs appear so users can see, edit and build on them in real time[1][3]
- It remained at ASL-2; Anthropic gave it to the UK AI Safety Institute for pre-deployment testing, which shared the results with the US AI Safety Institute[1]
- Anthropic said it would complete the family with Claude 3.5 Haiku and Claude 3.5 Opus later in 2024[1]
Notable features
- Priced at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window; AWS calls it 80 percent cheaper than Claude 3 Opus[1][5]
- Runs at twice the speed of Claude 3 Opus[1][4]
- Anthropic's model card: 59.4% on GPQA Diamond, 88.7% on MMLU, 92.0% on HumanEval and 71.1% on MATH, against Claude 3 Opus's 50.4%, 86.8%, 84.9% and 60.1%[2]
- It solved 64% of problems in an internal agentic coding evaluation - fixing a bug or adding a feature to an open-source codebase - against 38% for Claude 3 Opus[1][2]
- Anthropic's strongest vision model yet: better at reading charts and graphs and transcribing text from imperfect images, scoring 68.3% on MMMU[1][2]
Benchmarks[1]
| Benchmark | Claude 3.5 Sonnet | Claude 3 Opus | GPT-4o | Gemini 1.5 Pro | Llama-400b (early snapshot) |
|---|---|---|---|---|---|
| Graduate level reasoning (GPQA, Diamond) | 59.4%* (0-shot CoT) | 50.4% (0-shot CoT) | 53.6% (0-shot CoT) | — | — |
| Undergraduate level knowledge (MMLU), 5-shot | 88.7%** (5-shot) | 86.8% (5-shot) | — | 85.9% (5-shot) | 86.1% (5-shot) |
| Undergraduate level knowledge (MMLU), 0-shot CoT | 88.3% (0-shot CoT) | 85.7% (0-shot CoT) | 88.7% (0-shot CoT) | — | — |
| Code (HumanEval) | 92.0% (0-shot) | 84.9% (0-shot) | 90.2% (0-shot) | 84.1% (0-shot) | 84.1% (0-shot) |
| Multilingual math (MGSM) | 91.6% (0-shot CoT) | 90.7% (0-shot CoT) | 90.5% (0-shot CoT) | 87.5% (8-shot) | — |
| Reasoning over text (DROP, F1 score) | 87.1 (3-shot) | 83.1 (3-shot) | 83.4 (3-shot) | 74.9 (variable shots) | 83.5 (3-shot, pre-trained model) |
| Mixed evaluations (BIG-Bench-Hard) | 93.1% (3-shot CoT) | 86.8% (3-shot CoT) | — | 89.2% (3-shot CoT) | 85.3% (3-shot CoT, pre-trained model) |
| Math problem-solving (MATH) | 71.1% (0-shot CoT) | 60.1% (0-shot CoT) | 76.6% (0-shot CoT) | 67.7% (4-shot) | 57.8% (4-shot CoT) |
| Grade school math (GSM8K) | 96.4% (0-shot CoT) | 95.0% (0-shot CoT) | — | 90.8% (11-shot) | 94.1% (8-shot CoT) |
* Claude 3.5 Sonnet scores 67.2% on 5-shot CoT GPQA with maj@32. ** Claude 3.5 Sonnet scores 90.4% on MMLU with 5-shot CoT prompting.
References 585% CONFIDENCEOverall confidence: 85%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-3-5-sonnetanthropic.com· Posted Sep 28, 2026· Starts Jun 20, 2024· 32% of score
Anthropic[2]: "Today, we're launching Claude 3.5 Sonnet" and it "is now available for free on Claude.ai"; the page header now shows Jun 21, 2024, but TechCrunch[4] and AWS both published the launch on 20 June 2024 and the model ID is claude-3-5-sonnet-20240620.
- [2]90%Claude 3.5 Sonnet Model Card Addendumwww-cdn.anthropic.com· Added Sep 28, 2026· 32% of score
Anthropic's[1] model card addendum: 59.4% on GPQA Diamond, 88.7% on MMLU (5-shot), 92.0% on HumanEval, 71.1% on MATH and 68.3% on MMMU, against Claude 3 Opus's 50.4%, 86.8%, 84.9%, 60.1% and 59.4%; 64% on the internal agentic coding evaluation against Opus's 38%.
- [3]75%Claude (AI) - Wikipediaen.wikipedia.org· Added Sep 28, 2026· 32% of score
"On June 20, 2024, Anthropic[1][2] released Claude 3.5 Sonnet, which, according to the company's own benchmarks, performed better than the larger Claude 3 Opus", alongside Artifacts, which previews code, SVG graphics or websites in a separate window.
- [4]92%Anthropic claims its latest model is best-in-classtechcrunch.com· Published Jun 20, 2024· Starts Jun 20, 2024 ✓· 1% of score
TechCrunch, published 2024-06-20T14:00Z: Anthropic[1][2] "is releasing a powerful new generative AI model called Claude 3.5 Sonnet", about twice the speed of Claude 3 Opus, alongside Artifacts. It dates the launch 20 June, where Anthropic's own header now shows the 21 June UTC timestamp.
- [5]90%Anthropic's Claude 3.5 Sonnet model now available in Amazon Bedrock: Even more intelligence than Claude 3 Opus at one-fifth the costaws.amazon.com· Published Jun 20, 2024· Starts Jun 20, 2024· 1% of score
AWS News Blog, 20 JUN 2024: "Today, Anthropic[1][2] introduced Claude 3.5 Sonnet ... now available in Amazon Bedrock"; 80 percent cheaper than Opus and about 10 percent better than Claude 3 Opus on most vision benchmarks.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.