La prima voce è sempre la fonte del pin. L’affidabilità complessiva è una media ponderata di quanto ogni fonte conferma gli orari di inizio e fine indicati sopra; una fonte conta la metà per ogni 180 giorni in più rispetto alla più recente.
Le note di rilascio di Claude[2][3]: «November 4th, 2024 Claude Haiku 3.5 is now available on the Claude API as a text-only model»; TechCrunch[4] (10:24 PST, 4 novembre 2024) riporta «Anthropic's[6] newest AI model has arrived» e AWS ha pubblicato la disponibilità su Bedrock lo stesso giorno. Il post del 22 ottobre aveva detto che «sarà rilasciato entro la fine del mese».
Anthropic's[1][6] own dated changelog: "November 4th, 2024 Claude[3] Haiku 3.5 is now available on the Claude API as a text-only model", and on 24 February 2025 "We've added vision support to Claude Haiku 3.5". The model card addendum that is the source carries no release date, so this entry dates the pin.
Anthropic[1][6] notified developers on 19 December 2025 that Claude[2] Haiku 3.5 (claude-3-5-haiku-20241022) would be retired; it was retired on 19 February 2026, with claude-haiku-4-5-20251001 as the replacement.
TechCrunch (Kyle Wiggers, 10:24 AM PST, 4 November 2024): "Anthropic's[1][6] newest AI model has arrived"; it costs $1 per million input and $5 per million output tokens against Claude[2][3] 3 Haiku's $0.25 and $1.25, "a 4x hike", quoting Anthropic that "we've increased pricing for Claude 3.5 Haiku to reflect its increase in intelligence", and it launches without image analysis.
AWS What's New, "Posted on: Nov 4, 2024": Claude[2][3] 3.5 Haiku is now available in Amazon Bedrock and "surpasses even Claude 3 Opus, the largest model in Anthropic's[1][6] previous generation, on many intelligence benchmarks—including coding".
Manca qualcosa o c’è un errore? Dillo con parole tue: un link che conferma questo pin, una data di inizio o fine diversa e perché, o un fatto che manca o è sbagliato. L’IA lo confronta con le fonti del pin, ne cerca di migliori e aggiunge qualsiasi pagina ti dia ragione. Le fonti del pin contano comunque di più. Anche un’immagine che mostra altro, o lo mostra male, viene esaminata e spostata in basso o sostituita.
Titolo originale: Claude 3.5 Haiku Released
| Benchmark | Claude 3.5 Sonnet (nuovo) | Claude 3.5 Haiku | Claude 3.5 Sonnet | GPT-4o* | GPT-4o mini* | Gemini 1.5 Pro | Gemini 1.5 Flash |
|---|---|---|---|---|---|---|---|
| GPQA (Diamond) (ragionamento di livello post-laurea) | 65.0% 0-shot CoT | 41.6% 0-shot CoT | 59.4% 0-shot CoT | 53.6% 0-shot CoT | 40.2% 0-shot CoT | 59.1% 0-shot CoT | 51.0% 0-shot CoT |
| MMLU Pro (conoscenze di livello universitario) | 78.0% 0-shot CoT | 65.0% 0-shot CoT | 75.1% 0-shot CoT | — | — | 75.8% 0-shot CoT | 67.3% 0-shot CoT |
| HumanEval (codice) | 93.7% 0-shot | 88.1% 0-shot | 92.0% 0-shot | 90.2% 0-shot | 87.2% 0-shot | — | — |
| MATH (risoluzione di problemi matematici) | 78.3% 0-shot CoT | 69.2% 0-shot CoT | 71.1% 0-shot CoT | 76.6% 0-shot CoT | 70.2% 0-shot CoT | 86.5% 4-shot CoT | 77.9% 4-shot CoT |
| AIME 2024 (competizione di matematica delle superiori) | 16.0% 0-shot CoT | 5.3% 0-shot CoT | 9.6% 0-shot CoT | 9.3% 0-shot CoT | — | — | — |
| MMMU (domande e risposte visive) | 70.4% 0-shot CoT | — | 68.3% 0-shot CoT | 69.1% 0-shot CoT | 59.4% 0-shot CoT | 65.9% 0-shot CoT | 62.3% 0-shot CoT |
| SWE-bench Verified (programmazione con agenti) | 49.0% | 40.6% | 33.4% | — | — | — | — |
| TAU-bench (uso di strumenti con agenti) | Retail 69.2%, Compagnie aeree 46.0% | Retail 51.0%, Compagnie aeree 22.8% | Retail 62.6%, Compagnie aeree 36.0% | — | — | — | — |
* Le tabelle di valutazione di Anthropic escludono la famiglia di modelli o1 di OpenAI, che dipende da un lungo tempo di calcolo prima della risposta e rende difficili i confronti.