Originaltitel: Claude 3 Opus Released
| Benchmark | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| MMLU (Wissen auf Bachelor-Niveau) | 86,8 % 5-shot | 79,0 % 5-shot | 75,2 % 5-shot | 86,4 % 5-shot | 70,0 % 5-shot | 83,7 % 5-shot | 71,8 % 5-shot |
| GPQA, Diamond (Schlussfolgern auf Graduiertenniveau) | 50,4 % 0-shot CoT | 40,4 % 0-shot CoT | 33,3 % 0-shot CoT | 35,7 % 0-shot CoT | 28,1 % 0-shot CoT | — | — |
| GSM8K (Grundschulmathematik) | 95,0 % 0-shot CoT | 92,3 % 0-shot CoT | 88,9 % 0-shot CoT | 92,0 % 5-shot CoT | 57,1 % 5-shot | 94,4 % Maj1@32 | 86,5 % Maj1@32 |
| MATH (mathematisches Problemlösen) | 60,1 % 0-shot CoT | 43,1 % 0-shot CoT | 38,9 % 0-shot CoT | 52,9 % 4-shot | 34,1 % 4-shot | 53,2 % 4-shot | 32,6 % 4-shot |
| MGSM (mehrsprachige Mathematik) | 90,7 % 0-shot | 83,5 % 0-shot | 75,1 % 0-shot | 74,5 % 8-shot | — | 79,0 % 8-shot | 63,5 % 8-shot |
| HumanEval (Code) | 84,9 % 0-shot | 73,0 % 0-shot | 75,9 % 0-shot | 67,0 % 0-shot | 48,1 % 0-shot | 74,4 % 0-shot | 67,7 % 0-shot |
| DROP, F1-Wert (Schlussfolgern über Text) | 83,1 3-shot | 78,9 3-shot | 78,4 3-shot | 80,9 3-shot | 64,1 3-shot | 82,4 variable Shot-Anzahl | 74,1 variable Shot-Anzahl |
| BIG-Bench-Hard (gemischte Bewertungen) | 86,8 % 3-shot CoT | 82,9 % 3-shot CoT | 73,7 % 3-shot CoT | 83,1 % 3-shot CoT | 66,6 % 3-shot CoT | 83,6 % 3-shot CoT | 75,0 % 3-shot CoT |
| ARC-Challenge (Wissensfragen) | 96,4 % 25-shot | 93,2 % 25-shot | 89,2 % 25-shot | 96,3 % 25-shot | 85,2 % 25-shot | — | — |
| HellaSwag (Allgemeinwissen) | 95,4 % 10-shot | 89,0 % 10-shot | 85,9 % 10-shot | 95,3 % 10-shot | 85,5 % 10-shot | 87,8 % 10-shot | 84,7 % 10-shot |
Der erste Eintrag ist immer die Quelle des Pins. Die Gesamtsicherheit ist ein gewichteter Durchschnitt, wie fest jeder Beleg die oben verwendeten Start- und Endzeiten stützt; ein Beleg zählt für je 180 Tage, die er älter ist als der neueste, nur halb so viel.
Anthropics Beitrag ist datiert „Mar 4, 2024“ und sagt, „Opus and Sonnet are now available to use in claude[3].ai and the Claude API which is now generally available in 159 countries“; CNBC[4] und TechCrunch[5] berichten vom Launch am selben Montag.
Anthropic's[1] Claude[3] 3 model card (the link resolves to the PDF): the Claude 3 models' knowledge cutoff is August 2023 and they are offered through the Claude API, Amazon Bedrock and Google Vertex AI; it carries the full evaluations behind the launch post's table.
Anthropic's[1][2] deprecation history: on 30 June 2025 it notified developers of Claude Opus 3's retirement, and claude-3-opus-20240229 was retired on 5 January 2026 with claude-opus-4-8 as the recommended replacement.
CNBC's same-day report (published 2024-03-04T14:00Z): "Anthropic[1][2] on Monday debuted Claude[3] 3"; Opus outperformed GPT-4 and Gemini Ultra on benchmark tests, it is Anthropic's first multimodal model, Opus and Sonnet are available in 159 countries, and Anthropic's backers include Google, Salesforce and Amazon after about $7.3 billion in funding deals over the past year.
TechCrunch (Kyle Wiggers, 11:50 AM PST, 4 March 2024): Claude[3] 3 is Anthropic's[1][2] first multimodal model, can analyze up to 20 images in one request, starts with a 200,000-token context window (1 million for select customers), answers from data before August 2023, and Opus costs $15 per million input and $75 per million output tokens.
Fehlt etwas oder stimmt etwas nicht? Sag es in deinen eigenen Worten: ein Link, der diesen Pin belegt, ein anderes Start- oder Enddatum und warum, oder eine Angabe, die fehlt oder falsch ist. Die KI prüft es an den Quellen dieses Pins, sucht nach besseren und ergänzt jede Seite, die dich bestätigt. Die eigenen Quellen des Pins zählen weiterhin am meisten. Auch Bilder sieht sich die KI an: Eines, das etwas anderes oder die Sache schlecht zeigt, wird nach hinten gestellt oder ersetzt.