- 시작
- 2024년 3월 4일신뢰도 90%출처 제공
Claude 3 Opus 출시
원문에서 자동 번역됨
원본 제목: Claude 3 Opus Released
- Anthropic는 2024년 3월 4일, “역량이 오름차순인” Haiku, Sonnet, Opus로 구성된 Claude 3 패밀리를 발표했다. 같은 날 Opus와 Sonnet이 claude.ai와 Claude API에서 이용 가능해졌으며, 현재는 159개국에서 정식 서비스되고 있다.[1][4]
- claude.ai에서는 Opus가 Claude Pro 구독자에게, Sonnet이 무료 이용자에게 제공됐다. Sonnet은 같은 날 Amazon Bedrock에 도입됐고 Google Cloud의 Vertex AI에서는 비공개 프리뷰로 제공되며, “Opus와 Haiku도 두 곳 모두에 곧 도입될 예정”이었다.[1]
- 레드티밍 결과 “현재로서는 치명적 위험 가능성이 미미하다”고 판단돼, Anthropic는 이 모델을 AI Safety Level 2(ASL-2) 등급으로 유지했다.[1]
- CNBC는 Anthropic의 투자자로 Google, Salesforce, Amazon이 있으며, 이전 1년간 총 약 73억 달러 규모의 투자 계약이 있었다고 전했다.[4]
- Claude 3 Opus는 2025년 6월 30일 지원 종료(deprecated)됐고 2026년 1월 5일 Claude API에서 완전히 퇴역해 Claude Opus 4.8로 대체됐다.[3]
주요 특징
- Claude 3 Opus는 “매우 복잡한 작업에서 시장 최고 수준의 성능을 내는, 우리의 가장 지능적인 모델”로, 입력 토큰 100만 개당 15달러, 출력 토큰 100만 개당 75달러로 가격이 책정됐다.[1][5]
- Anthropic가 발표한 벤치마크: MMLU 86.8%, GPQA Diamond 50.4%, GSM8K 95.0%, HumanEval 84.9%로, GPT-4의 86.4%, 35.7%, 92.0%, 67.0%를 앞섰다.[1]
- 출시 당시 컨텍스트 윈도우는 20만 토큰이며, “특정 용도”로는 100만 토큰이 넘는 입력도 이용할 수 있었다. Needle In A Haystack 회상 테스트에서는 99%를 넘는 정확도를 기록했고, 때로는 삽입된 문장이 인위적으로 끼워 넣어졌다는 것을 알아채기도 했다.[1][5]
- Claude 최초의 비전 모델로, 사진·차트·그래프·기술 도표를 읽어내며 한 번의 요청에 이미지를 최대 20장까지 처리할 수 있지만 사람을 식별하지는 않는다.[1][5]
- Opus는 이전 Claude 모델들보다 무해한 프롬프트를 거부할 가능성이 “훨씬 낮으며”, 어려운 개방형 사실 질문에 대한 정확도는 Claude 2.1의 두 배다. 지식은 2023년 8월까지다.[1][2]
벤치마크[1]
| 벤치마크 | Claude 3 Opus | Claude 3 Sonnet | Claude 3 Haiku | GPT-4 | GPT-3.5 | Gemini 1.0 Ultra | Gemini 1.0 Pro |
|---|---|---|---|---|---|---|---|
| MMLU (학부 수준 지식) | 86.8% 5 shot | 79.0% 5-shot | 75.2% 5-shot | 86.4% 5-shot | 70.0% 5-shot | 83.7% 5-shot | 71.8% 5-shot |
| GPQA, Diamond (대학원 수준 추론) | 50.4% 0-shot CoT | 40.4% 0-shot CoT | 33.3% 0-shot CoT | 35.7% 0-shot CoT | 28.1% 0-shot CoT | — | — |
| GSM8K (초등 수학) | 95.0% 0-shot CoT | 92.3% 0-shot CoT | 88.9% 0-shot CoT | 92.0% 5-shot CoT | 57.1% 5-shot | 94.4% Maj1@32 | 86.5% Maj1@32 |
| MATH (수학 문제 해결) | 60.1% 0-shot CoT | 43.1% 0-shot CoT | 38.9% 0-shot CoT | 52.9% 4-shot | 34.1% 4-shot | 53.2% 4-shot | 32.6% 4-shot |
| MGSM (다국어 수학) | 90.7% 0-shot | 83.5% 0-shot | 75.1% 0-shot | 74.5% 8-shot | — | 79.0% 8-shot | 63.5% 8-shot |
| HumanEval (코드) | 84.9% 0-shot | 73.0% 0-shot | 75.9% 0-shot | 67.0% 0-shot | 48.1% 0-shot | 74.4% 0-shot | 67.7% 0-shot |
| DROP, F1 score (텍스트 기반 추론) | 83.1 3-shot | 78.9 3-shot | 78.4 3-shot | 80.9 3-shot | 64.1 3-shot | 82.4 Variable shots | 74.1 Variable shots |
| BIG-Bench-Hard (혼합 평가) | 86.8% 3-shot CoT | 82.9% 3-shot CoT | 73.7% 3-shot CoT | 83.1% 3-shot CoT | 66.6% 3-shot CoT | 83.6% 3-shot CoT | 75.0% 3-shot CoT |
| ARC-Challenge (지식 질의응답) | 96.4% 25-shot | 93.2% 25-shot | 89.2% 25-shot | 96.3% 25-shot | 85.2% 25-shot | — | — |
| HellaSwag (상식) | 95.4% 10-shot | 89.0% 10-shot | 85.9% 10-shot | 95.3% 10-shot | 85.5% 10-shot | 87.8% 10-shot | 84.7% 10-shot |
참고자료 5신뢰도 90%종합 신뢰도: 90%핀의 출처와 참고 자료가 날짜를 얼마나 뒷받침하는지.출처를 포함한 자료 5건이 핀의 시작·종료 시각을 얼마나 강하게 뒷받침하는지에 대한 가중 평균. 자료는 최신 자료보다 180일 오래될 때마다 가중치가 절반이 됩니다신뢰도 75% 이상 핀 모두 보기
첫 번째 항목은 항상 핀의 출처입니다. 전체 신뢰도는 각 자료가 위의 시작·종료 시각을 얼마나 확실히 뒷받침하는지에 대한 가중 평균이며, 자료는 최신 자료보다 180일 오래될 때마다 가중치가 절반이 됩니다.
- [1]90%anthropic.com/news/claude-3-familyanthropic.com· 게시 2026년 9월 28일· 시작 2024년 3월 4일 ✓· 점수의 33%
Anthropic[2]의 게시물은 “Mar 4, 2024”로 날짜가 표시돼 있으며 “Opus와 Sonnet은 이제 claude[3].ai와 Claude API에서 이용할 수 있으며, 159개국에서 정식 서비스된다”고 밝혔다. CNBC[4]와 TechCrunch[5]도 같은 월요일에 출시를 보도했다.
- [2]90%The Claude 3 Model Family: Opus, Sonnet, Haiku (model card)anthropic.com· 추가 2026년 9월 28일· 점수의 33%
Anthropic's[1] Claude[3] 3 model card (the link resolves to the PDF): the Claude 3 models' knowledge cutoff is August 2023 and they are offered through the Claude API, Amazon Bedrock and Google Vertex AI; it carries the full evaluations behind the launch post's table.
- [3]90%Model deprecations - Claude Platform Docsplatform.claude.com· 추가 2026년 9월 28일· 점수의 33%
Anthropic's[1][2] deprecation history: on 30 June 2025 it notified developers of Claude Opus 3's retirement, and claude-3-opus-20240229 was retired on 5 January 2026 with claude-opus-4-8 as the recommended replacement.
- [4]85%Anthropic, backed by Amazon and Google, debuts its most powerful chatbot yetcnbc.com· 게재 2024년 3월 4일· 시작 2024년 3월 4일· 점수의 1%
CNBC's same-day report (published 2024-03-04T14:00Z): "Anthropic[1][2] on Monday debuted Claude[3] 3"; Opus outperformed GPT-4 and Gemini Ultra on benchmark tests, it is Anthropic's first multimodal model, Opus and Sonnet are available in 159 countries, and Anthropic's backers include Google, Salesforce and Amazon after about $7.3 billion in funding deals over the past year.
- [5]85%Anthropic claims its new AI chatbot models beat OpenAI's GPT-4techcrunch.com· 게재 2024년 3월 4일· 시작 2024년 3월 4일· 점수의 1%
TechCrunch (Kyle Wiggers, 11:50 AM PST, 4 March 2024): Claude[3] 3 is Anthropic's[1][2] first multimodal model, can analyze up to 20 images in one request, starts with a 200,000-token context window (1 million for select customers), answers from data before August 2023, and Opus costs $15 per million input and $75 per million output tokens.
수정 제안
빠지거나 잘못된 점이 있나요? 이 핀을 뒷받침하는 링크, 다른 시작일이나 종료일과 그 이유, 빠지거나 잘못된 사실을 자신의 말로 알려주세요. AI가 이 핀의 출처와 대조하고 더 나은 자료를 찾아, 뒷받침하는 페이지를 추가합니다. 핀 자체의 출처가 여전히 가장 중요하게 반영됩니다. AI는 이미지도 확인합니다. 다른 것을 보여주거나 제대로 보여주지 못하는 이미지는 뒤로 옮기거나 교체합니다.