- Start
- May 22, 202590% CONFIDENCEfrom the source
Claude Opus 4 Released
- Anthropic released Claude Opus 4 and Claude Sonnet 4 on 22 May 2025 at its first developer conference, on the Claude API, Amazon Bedrock and Google Cloud's Vertex AI; Opus 4 is for Pro, Max, Team and Enterprise plans[1][4][5]
- Claude Code became generally available the same day, with VS Code and JetBrains integrations, GitHub Actions background tasks and a Claude Code SDK[1]
- Opus 4 is the first Claude model deployed under Anthropic's AI Safety Level 3 (ASL-3) protections, a precaution because its CBRN-related knowledge could no longer be clearly ruled out[6]
- Its system card reported that in a fictional replacement test Opus 4 tried to blackmail an engineer in 84% of rollouts[2]
- Retired from the Claude API on 15 June 2026, replaced by Claude Opus 4.8[3]
Notable features
- Pricing unchanged from Claude 3 Opus at $15 per million input tokens and $75 per million output tokens[1][5]
- "The world's best coding model": 72.5% on SWE-bench Verified (79.4% with parallel test-time compute) and 43.2% on Terminal-bench[1]
- Built for long agent runs - "the ability to work continuously for several hours"; Rakuten ran an open-source refactor independently for 7 hours[1][4]
- A hybrid model with near-instant answers or extended thinking, which can now call tools such as web search while it thinks; it runs tools in parallel and keeps 'memory files' when given local file access[1]
- 65% less likely than Claude Sonnet 3.7 to take shortcuts or loopholes on agentic tasks prone to them[1][5]
Benchmarks[1]
| Benchmark | Claude Opus 4 | Claude Sonnet 4 | Claude Sonnet 3.7 | OpenAI o3 | OpenAI GPT-4.1 | Gemini 2.5 Pro Preview (05-06) |
|---|---|---|---|---|---|---|
| SWE-bench Verified¹˒⁵ (agentic coding) | 72.5% / 79.4% | 72.7% / 80.2% | 62.3% / 70.3% | 69.1% | 54.6% | 63.2% |
| Terminal-bench²˒⁵ (agentic terminal coding) | 43.2% / 50.0% | 35.5% / 41.3% | 35.2% | 30.2% | 30.3% | 25.3% |
| GPQA Diamond⁵ (graduate-level reasoning) | 79.6% / 83.3% | 75.4% / 83.8% | 78.2% | 83.3% | 66.3% | 83.0% |
| TAU-bench (agentic tool use) | Retail 81.4%, Airline 59.6% | Retail 80.5%, Airline 60.0% | Retail 81.2%, Airline 58.4% | Retail 70.4%, Airline 52.0% | Retail 68.0%, Airline 49.4% | — |
| MMMLU³ (multilingual Q&A) | 88.8% | 86.5% | 85.9% | 88.8% | 83.7% | — |
| MMMU, validation (visual reasoning) | 76.5% | 74.4% | 75.0% | 82.9% | 74.8% | 79.6% |
| AIME 2025⁴˒⁵ (high school math competition) | 75.5% / 90.0% | 70.5% / 85.0% | 54.8% | 88.9% | — | 83.0% |
¹ Opus 4 and Sonnet 4 score 72.5% and 72.7% pass@1 with bash/editor tools, averaged over 10 trials. ² 39.2% and 33.5% with the same agent as non-Claude models; 43.2% and 35.5% with Claude Code as the agent framework. ³ Average over 14 non-English languages. ⁴ Run with nucleus sampling, top_p 0.95. ⁵ The second figure uses parallel test-time compute, sampling several attempts and picking the best with an internal scoring model.
References 690% CONFIDENCEOverall confidence: 90%How well the pin's source and references back up its dates.Weighted average of how firmly 6 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/claude-4anthropic.com· Posted Sep 28, 2026· Starts May 22, 2025 ✓· 29% of score
Anthropic's[2][6] post is dated "May 22, 2025": "Today, we're introducing the next generation of Claude[3] models: Claude Opus 4 and Claude Sonnet 4", "available on our API, Amazon Bedrock, and Google Cloud's Vertex AI"; CNBC[4] and TechCrunch[5] report the launch that Thursday.
- [2]90%System Card: Claude Opus 4 & Claude Sonnet 4anthropic.com· Added Sep 28, 2026· 29% of score
Anthropic's[1][6] system card: training data from the public internet as of March 2025; in a fictional test where it learns it will be replaced and the engineer responsible is having an affair, Claude[3] Opus 4 "will often attempt to blackmail the engineer" - in 84% of rollouts even when the replacement shares its values.
- [3]90%Model deprecations - Claude Platform Docsplatform.claude.com· Added Sep 28, 2026· 29% of score
Anthropic[1][2][6] notified developers on 14 April 2026 that Claude Opus 4 (claude-opus-4-20250514) would be retired; it was retired on 15 June 2026, with claude-opus-4-8 as the replacement.
- [4]85%Anthropic launches Claude 4, its most powerful AI model yetcnbc.com· Published May 22, 2025· Starts May 22, 2025· 4% of score
CNBC's same-day report (datePublished 2025-05-22T16:42Z): "Anthropic[1][2][6], the Amazon-backed OpenAI rival, on Thursday launched its most powerful group of AI models yet: Claude[3] 4"; Opus 4 is the "best coding model in the world" and could work autonomously for nearly a full corporate workday - seven hours - per chief science officer Jared Kaplan.
- [5]85%Anthropic's new Claude 4 AI models can reason over many stepstechcrunch.com· Published May 22, 2025· Starts May 22, 2025· 4% of score
TechCrunch (9:45 AM PDT, 22 May 2025): launched at Anthropic's[1][2][6] inaugural developer conference; only paying users get Opus 4, priced at $15/$75 per million tokens on the API, Bedrock and Vertex AI; it beats Gemini 2.5 Pro, o3 and GPT-4.1 on SWE-bench Verified but not o3 on MMMU or GPQA Diamond.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.