- Start
- Oct 22, 202490% CONFIDENCEfrom the source
Claude 3.5 Sonnet Upgrade Released
- Anthropic announced "an upgraded Claude 3.5 Sonnet" with "across-the-board improvements over its predecessor, with particularly significant gains in coding", available to all users that day, and the new Claude 3.5 Haiku for later in the month[1]
- Computer use arrived in public beta: Claude looks at the screen, moves a cursor, clicks buttons and types, and Anthropic calls 3.5 Sonnet "the first frontier AI model to offer computer use in public beta" - "still experimental—at times cumbersome and error-prone"[1][4]
- To steer the cursor Claude counts the pixels it needs to move on each screenshot; Anthropic trained the skill on simple software such as a calculator and a text editor[3]
- Developers could use it that day on the Anthropic API, Amazon Bedrock and Google Cloud's Vertex AI; Asana, Canva, Cognition, DoorDash, Replit and The Browser Company were early users, and Amazon had early access[1][5]
- The US and UK AI Safety Institutes tested it jointly before release; Anthropic kept it at ASL-2 and added classifiers that spot computer use and whether harm is happening[1]
Notable features
- Same price and speed as the June 3.5 Sonnet[1]
- SWE-bench Verified rises from 33.4% to 49.0%, above every publicly available model including OpenAI's o1-preview[1]
- TAU-bench agentic tool use rises from 62.6% to 69.2% in retail and from 36.0% to 46.0% in the airline domain[1]
- On OSWorld it scored 14.9% from screenshots alone against the next-best system's 7.8%, and 22.0% when allowed more steps[1][2]
- Scrolling, dragging and zooming were still hard for it, and Anthropic advised starting with low-risk tasks[1]
Benchmarks[1]
| Benchmark | Claude 3.5 Sonnet (new) | Claude 3.5 Haiku | Claude 3.5 Sonnet | GPT-4o* | GPT-4o mini* | Gemini 1.5 Pro | Gemini 1.5 Flash |
|---|---|---|---|---|---|---|---|
| Graduate level reasoning (GPQA Diamond) | 65.0% (0-shot CoT) | 41.6% (0-shot CoT) | 59.4% (0-shot CoT) | 53.6% (0-shot CoT) | 40.2% (0-shot CoT) | 59.1% (0-shot CoT) | 51.0% (0-shot CoT) |
| Undergraduate level knowledge (MMLU Pro) | 78.0% (0-shot CoT) | 65.0% (0-shot CoT) | 75.1% (0-shot CoT) | — | — | 75.8% (0-shot CoT) | 67.3% (0-shot CoT) |
| Code (HumanEval) | 93.7% (0-shot) | 88.1% (0-shot) | 92.0% (0-shot) | 90.2% (0-shot) | 87.2% (0-shot) | — | — |
| Math problem-solving (MATH) | 78.3% (0-shot CoT) | 69.2% (0-shot CoT) | 71.1% (0-shot CoT) | 76.6% (0-shot CoT) | 70.2% (0-shot CoT) | 86.5% (4-shot CoT) | 77.9% (4-shot CoT) |
| High school math competition (AIME 2024) | 16.0% (0-shot CoT) | 5.3% (0-shot CoT) | 9.6% (0-shot CoT) | 9.3% (0-shot CoT) | — | — | — |
| Visual Q/A (MMMU) | 70.4% (0-shot CoT) | — | 68.3% (0-shot CoT) | 69.1% (0-shot CoT) | 59.4% (0-shot CoT) | 65.9% (0-shot CoT) | 62.3% (0-shot CoT) |
| Agentic coding (SWE-bench Verified) | 49.0% | 40.6% | 33.4% | — | — | — | — |
| Agentic tool use (TAU-bench), retail | 69.2% | 51.0% | 62.6% | — | — | — | — |
| Agentic tool use (TAU-bench), airline | 46.0% | 22.8% | 36.0% | — | — | — | — |
* Anthropic's evaluation tables exclude OpenAI's o1 model family, as they depend on extensive pre-response computation time, unlike typical models.
References 590% CONFIDENCEOverall confidence: 90%How well the pin's source and references back up its dates.Weighted average of how firmly 5 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%anthropic.com/news/3-5-models-and-computer-useanthropic.com· Posted Sep 28, 2026· Starts Oct 22, 2024 ✓· 46% of score
The post is dated "Oct 22, 2024" and says "The upgraded Claude 3.5 Sonnet is now available for all users. Starting today, developers can build with the computer use beta"; TechCrunch[4] and CNBC[5] report the release "on Tuesday" the same day.
- [2]90%Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnetassets.anthropic.com· Added Sep 28, 2026· 46% of score
Anthropic's[1][3] model card addendum: the upgraded 3.5 Sonnet sets new state-of-the-art results in agentic coding (SWE-bench Verified), agentic tasks (TAU-bench) and computer use from screenshots (OSWorld), with a 14.9% average OSWorld success rate from screenshots alone.
- [3]90%Developing a computer use modelanthropic.com· Published Oct 22, 2024· 3% of score
Anthropic's[1][2] research post on the skill: Claude looks at screenshots and "counts how many pixels vertically or horizontally it needs to move a cursor in order to click in the correct place", and was trained on simple software such as a calculator and a text editor without internet access.
- [4]85%Anthropic's new AI model can control your PCtechcrunch.com· Published Oct 22, 2024· Starts Oct 22, 2024· 3% of score
TechCrunch, 2024-10-22: "Anthropic[1][2][3] on Tuesday released an upgraded version of its Claude 3.5 Sonnet model that can understand and interact with any desktop app" through a Computer Use API in open beta on the API, Bedrock and Vertex AI.
- [5]85%Amazon-backed Anthropic debuts AI agents that can do complex tasks, racing against OpenAI, Microsoft and Googlecnbc.com· Published Oct 22, 2024· Starts Oct 22, 2024· 3% of score
CNBC, 2024-10-22: chief science officer Jared Kaplan says it can do tasks with "tens or even hundreds of steps"; Amazon had early access and beta testers included Asana, Canva and Notion; released Tuesday in public beta for developers.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.