- Start
- Mar 4, 202495% CONFIDENCEfrom anthropic.com
Claude 3 Sonnet Released
- Anthropic announced the Claude 3 family - Haiku, Sonnet and Opus - on 4 March 2024, with Opus and Sonnet "now available to use in claude.ai and the Claude API", which became generally available in 159 countries; Haiku was to follow[3]
- Sonnet powers "the free experience on claude.ai", with Opus reserved for Claude Pro subscribers, and launched the same day in Amazon Bedrock and in private preview on Google Cloud's Vertex AI Model Garden[3][4]
- Anthropic pitched it as the model that "strikes the ideal balance between intelligence and speed—particularly for enterprise workloads", and for most workloads twice as fast as Claude 2 and Claude 2.1[3][4]
- The family stayed at AI Safety Level 2 (ASL-2) under Anthropic's Responsible Scaling Policy, and refuses harmless prompts near its guardrails far less often than earlier Claude models[3]
- When Anthropic retired Claude 3 Sonnet in July 2025, around 200 people gathered in San Francisco for a "funeral"[2]
Notable features
- Priced at $3 per million input tokens and $15 per million output tokens, with a 200K-token context window[3]
- The first Claude generation with vision: it reads photos, charts, graphs and technical diagrams, up to 20 images in one request, though Anthropic stopped it identifying people[1][5]
- Model card benchmarks: 79.0% on MMLU (5-shot), 40.4% on GPQA Diamond, 92.3% on GSM8K and 73.0% on HumanEval Python coding[1]
- Trained on Amazon Web Services and Google Cloud hardware with PyTorch, JAX and Triton; it answers from data before August 2023 and cannot search the web[1][5]
- Anthropic's suggested uses: retrieval over large knowledge bases, product recommendations and forecasting, code generation and parsing text from images[3]
Benchmarks[3]
| Benchmark | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| Undergraduate level knowledge (MMLU) | 86.8% (5-shot) | 79.0% (5-shot) | 75.2% (5-shot) | 86.4% (5-shot) | 70.0% (5-shot) | 83.7% (5-shot) | 71.8% (5-shot) |
| Graduate level reasoning (GPQA, Diamond) | 50.4% (0-shot CoT) | 40.4% (0-shot CoT) | 33.3% (0-shot CoT) | 35.7% (0-shot CoT) | 28.1% (0-shot CoT) | — | — |
| Grade school math (GSM8K) | 95.0% (0-shot CoT) | 92.3% (0-shot CoT) | 88.9% (0-shot CoT) | 92.0% (5-shot CoT) | 57.1% (5-shot) | 94.4% (Maj1@32) | 86.5% (Maj1@32) |
| Math problem-solving (MATH) | 60.1% (0-shot CoT) | 43.1% (0-shot CoT) | 38.9% (0-shot CoT) | 52.9% (4-shot) | 34.1% (4-shot) | 53.2% (4-shot) | 32.6% (4-shot) |
| Multilingual math (MGSM) | 90.7% (0-shot) | 83.5% (0-shot) | 75.1% (0-shot) | 74.5% (8-shot) | — | 79.0% (8-shot) | 63.5% (8-shot) |
| Code (HumanEval) | 84.9% (0-shot) | 73.0% (0-shot) | 75.9% (0-shot) | 67.0% (0-shot) | 48.1% (0-shot) | 74.4% (0-shot) | 67.7% (0-shot) |
| Reasoning over text (DROP, F1 score) | 83.1 (3-shot) | 78.9 (3-shot) | 78.4 (3-shot) | 80.9 (3-shot) | 64.1 (3-shot) | 82.4 (variable shots) | 74.1 (variable shots) |
| Mixed evaluations (BIG-Bench-Hard) | 86.8% (3-shot CoT) | 82.9% (3-shot CoT) | 73.7% (3-shot CoT) | 83.1% (3-shot CoT) | 66.6% (3-shot CoT) | 83.6% (3-shot CoT) | 75.0% (3-shot CoT) |
| Knowledge Q&A (ARC-Challenge) | 96.4% (25-shot) | 93.2% (25-shot) | 89.2% (25-shot) | 96.3% (25-shot) | 85.2% (25-shot) | — | — |
| Common knowledge (HellaSwag) | 95.4% (10-shot) | 89.0% (10-shot) | 85.9% (10-shot) | 95.3% (10-shot) | 85.5% (10-shot) | 87.8% (10-shot) | 84.7% (10-shot) |
References 583% CONFIDENCEOverall confidence: 83%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/claude-3-model-cardanthropic.com· Posted Sep 28, 2026· Starts Mar 4, 2024· 48% of score
The model card carries no launch date, so the date is the launch post's: "Introducing the next generation of Claude, Mar 4, 2024" - "Opus and Sonnet are now available to use in claude.ai and the Claude API"; AWS's post of 04 MAR 2024 announces "the availability of Anthropic's[3] Claude 3 Sonnet today in Amazon[4] Bedrock".
- [2]75%Claude (AI) - Wikipediaen.wikipedia.org· Added Sep 28, 2026· 48% of score
The Sonnet table lists Claude 3 Sonnet as released 4 March 2024 and discontinued; the article adds that when Anthropic[1][3] retired it in July 2025 around 200 people gathered in San Francisco for a "funeral".
- [3]95%Introducing the next generation of Claudeanthropic.com· Published Mar 4, 2024· Starts Mar 4, 2024 ✓· 1% of score
Anthropic's[1] launch post, dated "Mar 4, 2024": "Opus and Sonnet are now available to use in claude.ai and the Claude API which is now generally available in 159 countries", with Sonnet at $3/$15 per million tokens and a 200K context window, powering the free claude.ai. The model card that is the source carries no launch date, so this row dates the pin.
- [4]90%Anthropic's Claude 3 Sonnet foundation model is now available in Amazon Bedrockaws.amazon.com· Published Mar 4, 2024· Starts Mar 4, 2024· 1% of score
AWS News Blog post by Channy Yun, 04 MAR 2024: "We're also announcing the availability of Anthropic's[1][3] Claude 3 Sonnet today in Amazon Bedrock, with Claude 3 Opus and Claude 3 Haiku coming soon"; Sonnet is two times faster than Claude 2 and 2.1.
- [5]85%Anthropic claims its new AI chatbot models beat OpenAI's GPT-4techcrunch.com· Published Mar 4, 2024· Starts Mar 4, 2024· 1% of score
TechCrunch's same-day report: Opus and Sonnet "are available now on the web and via Anthropic's[1][3] dev console and API, Amazon's[4] Bedrock platform and Google's Vertex AI"; Claude 3 is Anthropic's first multimodal model, takes up to 20 images per request, cannot identify people and answers only from data before August 2023.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.