- Start
- Mar 4, 202490% CONFIDENCEfrom the source
Claude 3 Opus Released
- Anthropic announced the Claude 3 family - Haiku, Sonnet and Opus "in ascending order of capability" - on 4 March 2024, with Opus and Sonnet available the same day in claude.ai and the Claude API, now generally available in 159 countries[1][4]
- On claude.ai Opus went to Claude Pro subscribers while Sonnet powered the free experience; Sonnet reached Amazon Bedrock and a Google Cloud Vertex AI private preview that day, "with Opus and Haiku coming soon to both"[1]
- Anthropic kept the model at AI Safety Level 2 (ASL-2) after red-teaming found "negligible potential for catastrophic risk at this time"[1]
- CNBC notes Anthropic's backers include Google, Salesforce and Amazon, after funding deals totalling about $7.3 billion over the previous year[4]
- Claude 3 Opus was deprecated on 30 June 2025 and retired from the Claude API on 5 January 2026, replaced by Claude Opus 4.8[3]
Notable features
- Claude 3 Opus is "our most intelligent model, with best-in-market performance on highly complex tasks", priced at $15 per million input tokens and $75 per million output tokens[1][5]
- Benchmarks Anthropic reports: 86.8% on MMLU, 50.4% on GPQA Diamond, 95.0% on GSM8K and 84.9% on HumanEval, ahead of GPT-4's 86.4%, 35.7%, 92.0% and 67.0%[1]
- A 200K-token context window at launch, with inputs over 1 million tokens available "for specific use cases"; on Needle In A Haystack recall it passed 99% accuracy and at times noticed the planted sentence looked artificially inserted[1][5]
- Claude's first vision model: it reads photos, charts, graphs and technical diagrams, up to 20 images in one request, but will not identify people[1][5]
- Opus is "significantly less likely" than earlier Claude models to refuse harmless prompts, and doubles Claude 2.1's accuracy on hard open-ended factual questions; knowledge runs to August 2023[1][2]
Benchmarks[1]
| Benchmark | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| MMLU (undergraduate level knowledge) | 86.8% 5 shot | 79.0% 5-shot | 75.2% 5-shot | 86.4% 5-shot | 70.0% 5-shot | 83.7% 5-shot | 71.8% 5-shot |
| GPQA, Diamond (graduate level reasoning) | 50.4% 0-shot CoT | 40.4% 0-shot CoT | 33.3% 0-shot CoT | 35.7% 0-shot CoT | 28.1% 0-shot CoT | — | — |
| GSM8K (grade school math) | 95.0% 0-shot CoT | 92.3% 0-shot CoT | 88.9% 0-shot CoT | 92.0% 5-shot CoT | 57.1% 5-shot | 94.4% Maj1@32 | 86.5% Maj1@32 |
| MATH (math problem-solving) | 60.1% 0-shot CoT | 43.1% 0-shot CoT | 38.9% 0-shot CoT | 52.9% 4-shot | 34.1% 4-shot | 53.2% 4-shot | 32.6% 4-shot |
| MGSM (multilingual math) | 90.7% 0-shot | 83.5% 0-shot | 75.1% 0-shot | 74.5% 8-shot | — | 79.0% 8-shot | 63.5% 8-shot |
| HumanEval (code) | 84.9% 0-shot | 73.0% 0-shot | 75.9% 0-shot | 67.0% 0-shot | 48.1% 0-shot | 74.4% 0-shot | 67.7% 0-shot |
| DROP, F1 score (reasoning over text) | 83.1 3-shot | 78.9 3-shot | 78.4 3-shot | 80.9 3-shot | 64.1 3-shot | 82.4 Variable shots | 74.1 Variable shots |
| BIG-Bench-Hard (mixed evaluations) | 86.8% 3-shot CoT | 82.9% 3-shot CoT | 73.7% 3-shot CoT | 83.1% 3-shot CoT | 66.6% 3-shot CoT | 83.6% 3-shot CoT | 75.0% 3-shot CoT |
| ARC-Challenge (knowledge Q&A) | 96.4% 25-shot | 93.2% 25-shot | 89.2% 25-shot | 96.3% 25-shot | 85.2% 25-shot | — | — |
| HellaSwag (common knowledge) | 95.4% 10-shot | 89.0% 10-shot | 85.9% 10-shot | 95.3% 10-shot | 85.5% 10-shot | 87.8% 10-shot | 84.7% 10-shot |
References 590% CONFIDENCEOverall confidence: 90%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-3-familyanthropic.com· Posted Sep 28, 2026· Starts Mar 4, 2024 ✓· 33% of score
Anthropic's[2] post is dated "Mar 4, 2024" and says "Opus and Sonnet are now available to use in claude[3].ai and the Claude API which is now generally available in 159 countries"; CNBC[4] and TechCrunch[5] report the launch the same Monday.
- [2]90%The Claude 3 Model Family: Opus, Sonnet, Haiku (model card)anthropic.com· Added Sep 28, 2026· 33% of score
Anthropic's[1] Claude[3] 3 model card (the link resolves to the PDF): the Claude 3 models' knowledge cutoff is August 2023 and they are offered through the Claude API, Amazon Bedrock and Google Vertex AI; it carries the full evaluations behind the launch post's table.
- [3]90%Model deprecations - Claude Platform Docsplatform.claude.com· Added Sep 28, 2026· 33% of score
Anthropic's[1][2] deprecation history: on 30 June 2025 it notified developers of Claude Opus 3's retirement, and claude-3-opus-20240229 was retired on 5 January 2026 with claude-opus-4-8 as the recommended replacement.
- [4]85%Anthropic, backed by Amazon and Google, debuts its most powerful chatbot yetcnbc.com· Published Mar 4, 2024· Starts Mar 4, 2024· 1% of score
CNBC's same-day report (published 2024-03-04T14:00Z): "Anthropic[1][2] on Monday debuted Claude[3] 3"; Opus outperformed GPT-4 and Gemini Ultra on benchmark tests, it is Anthropic's first multimodal model, Opus and Sonnet are available in 159 countries, and Anthropic's backers include Google, Salesforce and Amazon after about $7.3 billion in funding deals over the past year.
- [5]85%Anthropic claims its new AI chatbot models beat OpenAI's GPT-4techcrunch.com· Published Mar 4, 2024· Starts Mar 4, 2024· 1% of score
TechCrunch (Kyle Wiggers, 11:50 AM PST, 4 March 2024): Claude[3] 3 is Anthropic's[1][2] first multimodal model, can analyze up to 20 images in one request, starts with a 200,000-token context window (1 million for select customers), answers from data before August 2023, and Opus costs $15 per million input and $75 per million output tokens.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.