Anthropic today announced it's releasing Claude Sonnet 5.5, a faster and more efficient update to its mid-tier, workhorse AI model.

In several of the company’s tests, Sonnet 5.5 comes strikingly close to the performance of its more expensive sibling, Opus 5.5, Anthropic's new flagship model released just last week — giving enterprises another alternative to paying for a top-tier frontier model.

Anthropic says Sonnet 5.5 generates output more than 30% faster than Claude Sonnet 5 and can reduce the total cost of completing a task by as much as 30%, primarily because it uses fewer tokens and fewer tool calls rather than because of a lower API sticker price. VentureBeat reached out to inquire exactly how Anthropic was measuring this speed bump and will update when we receive a response.

Meanwhile, Anthropic is keeping the application programming interface (API) price of Sonnet 5.5 the same as its immediate predecessor, Sonnet 5m, at $2/$10 per 1 million input/output tokens, respectively, alongside a discount of $0.20 per million cache-read tokens (sending the same information as input). Cache writes also cost less, at $2.50 per million tokens.

By comparison, Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with $5 cache writes.

In other words, Anthropic’s pitch is increasingly about the cost of accomplishing a job, not merely the price of processing an individual token.

Anthropic pushes more Opus capability into Sonnet

Anthropic describes Sonnet 5.5 as best suited to relatively well-defined everyday work: software debugging, coding, producing documents, building presentations and spreadsheets, and designing or refining interfaces. Opus 5.5 remains the company’s preferred option for ambiguous or open-ended work that requires sustained judgment.

But Anthropic’s own evaluation results suggest the gap between the two tiers has narrowed substantially in some workloads.

On GDPval-AA, which measures performance on tasks drawn from real-world occupations, Anthropic reports Sonnet 5.5 scoring 1844, almost identical to Opus 5.5’s 1846 and well ahead of Sonnet 5’s 1449. On AA-Briefcase, Sonnet 5.5 scores 1811 compared with 1822 for Opus 5.5.

The model also reaches 80.1% on OSWorld 2.1, Anthropic’s partial computer-use evaluation, compared with 81.8% for Opus 5.5 and 57% for Sonnet 5. On Chartography, a visual chart-recognition test, Sonnet 5.5 scores 61.6% versus 64.4% for Opus 5.5 and just 15.6% for Sonnet 5.

Anthropic says the model is particularly strong at coding. On Terminal-Bench 4.0, an agentic coding benchmark, it reports Sonnet 5.5 at 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 under the reported settings. On CursorBench 4.0, Sonnet 5.5 reaches 55.5%, narrowly behind Opus 5.5 at 57.8%.

Benchmarks should not be treated as direct proxies for every production workload, and Anthropic itself says Opus 5.5 remains clearly stronger at difficult, open-ended work.

Still, the economics could matter more than the leaderboard.

Anthropic says that on several evaluations, Sonnet 5.5 running at Low or Medium effort can surpass Sonnet 5’s best score at roughly one-tenth the cost per task. On FrontierCode, it says Sonnet 5.5 at High effort scores about 10 points higher than Sonnet 5 at the same setting while costing around one-fifteenth as much per task.

Claude Code and Anthropic’s consumer applications will default the model to Medium effort, while the Claude Platform defaults to High. Lower settings trade some reasoning depth for lower latency and token consumption; higher settings allow the model to spend longer checking and refining its work.

Customers report fewer steps, faster agents

Anthropic also supplied VentureBeat with early customer testing that offers a more practical picture of what those efficiency claims could mean inside enterprise workflows.

At Box, VP of AI Products Yashodha Bhavnani said Sonnet 5.5 rechecked source documents and caught errors its predecessor missed.

“Compared to the last model, Sonnet 5.5 was more accurate, 2.4x faster and used 12% fewer total tokens,” Bhavnani said.

Zendesk tested the model across hundreds of support cases spanning customer replies and escalation decisions. Director of AI Abhinay Kathuria said tickets were processed 20% faster and that the model made fewer incorrect decisions than Claude models Zendesk currently uses in production.

Slack saw a similar pattern. Principal Engineer Curtis Allen said Sonnet 5.5 outperformed Sonnet 5 on nearly all of the company’s offline Slackbot evaluations without any prompt changes, while requiring fewer steps and roughly 14% fewer output tokens.

For coding agents, the reduction in steps may prove especially important. Lovable co-founder and CTO Fabian Hedin said its evaluations found Sonnet 5.5 required about one-third fewer tool calls and roughly half as many shell executions to complete coding jobs.

And Base44 found that across 118 real application builds, Sonnet 5.5 reached results comparable to Opus 5 in an average of 3.6 iterations per build, versus 7.7 for Opus 5, while producing the fewest failed tool calls among the models it compared.

Those numbers underline an increasingly important consideration for enterprises deploying AI agents: token prices tell only part of the story. Every unnecessary tool call, failed action and retry can add latency, infrastructure cost and another opportunity for an autonomous workflow to go off course.

The price war around everyday frontier models is getting tighter

Sonnet 5.5 also arrives in an increasingly aggressive market for models positioned below the most expensive frontier tier.

OpenAI released GPT-6 Sol last week and currently lists its Standard API pricing at the same $2 per million input tokens and $10 per million output tokens as Sonnet 5.5, along with $0.20 cached input and $2.50 cache writes.

OpenAI positions Sol for complex coding and agentic workflows and gives it a 1.05-million-token context window.

Google is undercutting both companies on headline price with Gemini 3.8 Flash, which is currently offered at an introductory $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Google says standard pricing will rise to $1.50 and $7.50 respectively on January 1, 2027.

And recently, Chinese electronic and rapidly growing AI research lab Xiaomi released its frontier-class MiMo-V2.6-Pro and mid-tier Flash models under a permissive, enterprise-friendly MIT License, meaning enterprises and developers can download these models freely now and customize them to their liking, without paying Xiaomi a cent. Over the API, they remain among the more affordable models in the world.

Model

Input ($/1M)

Output ($/1M)

Total ($/1M)

Source

Muse Spark 1.2 / 1.3 Contributor

$0.10

$0.20

$0.30

Meta

MiMo-V2.6-Flash

$0.14

$0.28

$0.42

Xiaomi

GPT-6 Luna

$0.10

$0.50

$0.60

OpenAI

DeepSeek-V4.1-Flash — off-peak

$0.15

$0.60

$0.75

DeepSeek

MiMo-V2.6-Pro

$0.435

$0.87

$1.305

Xiaomi

MiniMax-M3

$0.30

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

$0.30

$1.20

$1.50

LongCat

DeepSeek-V4.1-Flash — peak hours

$0.30

$1.20

$1.50

DeepSeek

DeepSeek-V4-Pro — off-peak

$0.66

$1.98

$2.64

DeepSeek

LongCat-2.0 — standard

$0.75

$2.95

$3.70

LongCat

Gemini 3.7 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

Gemini 3.8 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

DeepSeek-V4-Pro — peak hours

$1.32

$3.96

$5.28

DeepSeek

Muse Spark 1.1 / 1.2 / 1.3

$1.25

$4.25

$5.50

Meta

GLM-5.3

$1.40

$4.40

$5.80

Z.AI

Grok 4.7 — ≤200K prompt tokens

$2.00

$6.00

$8.00

xAI

Qwen3.8-Max

$2.00

$6.00

$8.00

QwenCloud

Gemini 3.7 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

Gemini 3.8 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

GPT-6 Sol

$2.00

$10.00

$12.00

OpenAI

Claude Sonnet 5.5

$2.00

$10.00

$12.00

Anthropic

Grok 4.7 — >200K prompt tokens

$4.00

$12.00

$16.00

xAI

Grok 4.7 Fast (Cursor) — ≤256K input

$4.00

$12.00

$16.00

Cursor

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Kimi K3

$3.00

$15.00

$18.00

Moonshot AI

Claude Opus 5.5

$4.00

$20.00

$24.00

Anthropic

Grok 4.7 Fast (Cursor) — >256K input, up to 500K

$6.00

$18.00

$24.00

Cursor

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Sakana AI

Claude Opus 5.5 — Fast mode

$8.00

$40.00

$48.00

Anthropic

Claude Fable 5 / Claude Mythos 5

$10.00

$50.00

$60.00

Anthropic

Claude Fable 5.1 / Claude Mythos 5.1

$10.00

$50.00

$60.00

Anthropic

GPT-6 Astra — Standard mode

$10.00

$50.00

$60.00

OpenAI

GPT-6 Astra — Fast mode

$20.00

$100.00

$120.00

OpenAI

That means Anthropic is not trying to win this round simply by having the cheapest tokens. Instead, Sonnet 5.5 makes a broader efficiency argument: that a model which completes a job in fewer reasoning steps, tokens and tool calls can deliver a lower total operating cost even when its per-token pricing stays unchanged.

That argument also represents a change from the original Sonnet 5 launch. Anthropic introduced Sonnet 5 in June at $2 input and $10 output per million tokens as introductory pricing, before making those rates permanent in August rather than moving to the previously planned $3/$15 rate.

Stronger cyber capabilities bring stronger restrictions

The performance improvement comes with a notable safety change: Sonnet 5.5 is the first Sonnet model Anthropic is launching with cybersecurity safeguards modeled on those used for its highest-capability systems.

Anthropic says routine software development and vulnerability remediation should continue to work normally, but some higher-risk cybersecurity requests will automatically fall back to Sonnet 5. The company plans to expand its Cyber Verification Program so approved defenders can access more advanced capabilities in Sonnet 5.5, Opus 5.5 and its Mythos models.

Anthropic is also introducing safeguards intended to prevent what it calls "model distillation attacks," in which attackers attempt to reproduce a model’s capabilities by querying it at industrial scale. Critics note that many modern open LLMs are trained off data generated by others, and that Anthropic, having already ingested mass amounts of copyrighted data without express consent or payment from many authors and rights-holders, is not in a strong position ethically or morally to argue that deriving a new model off its own model outputs is somehow criminal or wrong.

Sonnet 5.5 is the first Sonnet model to ship with classifiers designed to block extraction of its reasoning, while Anthropic is expanding its “preserved thinking” system so reasoning cannot be detached from the account that created it. Biology safeguards remain the same as Sonnet 5.

Available across all major clouds

Claude Sonnet 5.5 will be available through Anthropic as well as Amazon Web Services, Google Cloud and Microsoft Azure, with zero-data-retention support. Developers can access it through the Claude Platform using the model identifier claude-sonnet-5-5.

Anthropic says Claude Haiku 5.5, aimed at high-volume and especially cost-sensitive workloads, will follow in the coming weeks.

For enterprise developers, however, Sonnet 5.5 may be the more consequential part of the family: not because Anthropic has dramatically lowered its API rates, but because the company is trying to make the middle tier perform enough like the premium tier that many workloads no longer need to climb higher.