AI Research Atlas

Claude 3.5 Sonnet (upgraded)

Anthropic · 22 October 2024

An upgraded Claude 3.5 Sonnet lifts SWE-bench Verified from 33.4% to 49.0% at unchanged price and speed.

Same name, new weights, with gains across the board and the largest in coding and tool use. SWE-bench Verified rises 33.4% to 49.0%, TAU-bench retail 62.6% to 69.2% and airline 36.0% to 46.0%. Price and speed unchanged. It is the model that first shipped with computer use.

Date
Tuesday, 22 October 2024
Lab
Anthropic
Kind
model
Access
closed API
Price
$3/$15 per M tokens in/out, unchanged

Figures

MeasureValueMeasured by
SWE-bench Verified49.0%
from 33.4% for the June 3.5 Sonnet
company
TAU-bench retail69.2%
from 62.6%
company
TAU-bench airline46.0%
from 36.0%
company

The naming is confusing. Both the June and October models are called Claude 3.5 Sonnet, and the October one is claude-3-5-sonnet-20241022 (the model ID is not on the page captured here).

Sources

  1. www.anthropic.com/news/3-5-models-and-computer-use

This record was checked against its sources on 6 October 2026. How we check

Related