Anthropic today released Claude Haiku 5.5, cutting token prices by 90% for requests below 100,000 tokens and targeting the repetitive work that can make enterprise AI expensive at scale: summarizing documents, classifying information, querying databases and handling smaller assignments for more capable agents.

The model starts at $0.10 per million input tokens and $0.50 per million output tokens, matching OpenAI’s GPT-6 Luna. Anthropic estimates that workloads cost approximately 75% less to run than on Haiku 4.5, once request sizes and changes in token consumption are considered.

The launch also brings lower Sonnet 5.5 caching charges and monthly API credits for Max and Team subscribers, broadening the announcement beyond a single inexpensive model.

A support role for enterprise agents

Anthropic positions Haiku as a supporting worker for Opus and Sonnet, which remain its recommended choices for demanding coding assignments. A larger model might assemble a financial presentation while Haiku retrieves the revenue figure needed for one slide, according to an example supplied by financial AI company Rogo.

“It's accurate enough that we'd trust it there and fast and cheap enough that we can run it a lot,” Alex Wang, who works in applied AI at Rogo, said in a statement provided by Anthropic.

The largest savings apply to the shortest requests

The pricing has an important boundary: above 100,000 tokens, Haiku 5.5 costs $0.50 per million input tokens and $2.50 per million output tokens. Those rates represent a 50% reduction from Haiku 4.5, compared with the 90% reduction below that threshold.

Charge, per million tokens

Haiku 5.5: below 100,000 tokens

Haiku 5.5: above 100,000 tokens

Haiku 4.5

Input

$0.10

$0.50

$1.00

Output

$0.50

$2.50

$5.00

Cache reads

$0.01

$0.05

$0.10

Cache writes

$0.125

$0.625

$1.25

Source: Anthropic’s launch materials. The materials do not specify how requests of exactly 100,000 tokens are billed or precisely which tokens determine the threshold.

Anthropic says approximately 90% of Haiku 4.5 requests fall into the shorter category. Its estimated 75% average workload savings also accounts for an updated tokenizer—the mechanism that divides content into billable units—which uses somewhat more tokens for equivalent work.

How Haiku compares with OpenAI, Google and Grok

The lower rates make Haiku competitive with other proprietary models, but do not establish it as the cheapest option across the market. OpenAI’s Luna matches all four of Haiku’s lower-tier rates, including caching.

Model

Input per million tokens

Output per million tokens

Pricing qualification

Anthropic Claude Haiku 5.5

$0.10

$0.50

Requests below 100,000 tokens

OpenAI GPT-6 Luna

$0.10

$0.50

Base rates; higher rates above 272,000 input tokens

Google Gemini 3.5 Flash-Lite

$0.30

$2.50

Standard pricing

Anthropic Claude Haiku 5.5

$0.50

$2.50

Requests above 100,000 tokens

Google Gemini 3.8 Flash

$0.75

$3.75

Promotional standard rates through Dec. 31, 2026

Grok 4.3

$1.25

$2.50

Below 200,000 prompt tokens

Grok 4.7

$2.00

$6.00

Below 200,000 prompt tokens

OpenAI GPT-6.1 Sol

$2.00

$10.00

Base standard rates

Anthropic Claude Sonnet 5.5

$2.00

$10.00

Listed launch-table rates

Sources: Anthropic’s supplied materials; official pricing for GPT-6 Luna, GPT-6.1 Sol, Gemini and Grok, checked October 7. Figures exclude caching, batch discounts, tools, regional premiums and negotiated rates. This compares prices, not equivalent capabilities.

The context thresholds matter: Haiku’s higher tier is five times its lower rate, while Luna’s surcharge begins above 272,000 input tokens. Token prices alone also cannot establish the cost of completing a job; token consumption, retries and accuracy affect that calculation.

Benchmark gains come with an effort-setting caveat

Anthropic’s benchmark table shows gains over Haiku 4.5 and leads over GPT-6 Luna on the selected evaluations below, while Sonnet 5.5 remains ahead. These are vendor-reported results, not independent verification.

Evaluation

Haiku 5.5

Haiku 4.5

GPT-6 Luna

Sonnet 5.5

GDPval-AA v2.1: knowledge work

1,620

735

1,437

1,840

AA-Briefcase v1.1: knowledge work

1,578

614

1,336

1,824

OSWorld 2.1, offline subset: computer operation

72.4%

15.7%

48.9%

83.9%

Terminal-Bench 4.0: agent-based coding

39.2%

0.0%

16.4%

70.6%

FrontierCode 1.1, main evaluation

46.4%

—

42.4%

52.1%*

Source: Anthropic. Sonnet’s FrontierCode result is labeled High effort; the dash indicates no result supplied. The first two rows are scores, not percentages.

Haiku 5.5 introduces adjustable effort levels, with medium as the default. That distinction matters for interpreting its results: Anthropic’s Terminal-Bench chart places the approximately 39% score at maximum effort, while medium scores approximately 20%.

Customers report faster results, but throughput remains unspecified

Anthropic calls Haiku its fastest model at standard speeds, while acknowledging that Opus in Fast Mode runs faster. It does not supply a tokens-per-second figure in the materials reviewed for this article.

Customer reports offer narrower evidence. Box’s VP of AI Products, Yashodha Bhavnani, reports an 11-point improvement over Haiku 4.5 with approximately half the latency, without identifying the scoring scale. Asana’s Aaron Vinh, a staff software engineer, reports task-completion latency falling more than 30% and inference per agent turn accelerating by as much as 2.5 times, compared with an unnamed model.

HubSpot’s Ze’ev Klapow, a distinguished software engineer, reports a 92.8% average across three runs of its CRM evaluation, the strongest result among the models it tested. These company-supplied testimonials do not establish comparable throughput across providers.

Sonnet price cuts and new subscription API credits

Alongside Haiku, Sonnet 5.5 cache reads are getting a cost cut from $0.20 to $0.10 per million tokens.

Anthropic estimates approximately 20% savings on typical agent workloads; actual savings depend on how much stored context an application reuses.

Monthly API credits roll out this week: $100 for Max 5x subscribers, $200 for Max 20x and up to $500 shared among a Team subscription’s users. They can be spent on any model through Anthropic’s platform. The supplied draft does not explain Team allocation or rollover conditions.

Cloud availability and deployment considerations

Anthropic says Haiku launches through its own platform, AWS, Google Cloud and Microsoft Azure, with the API identifier claude-haiku-5-5. Python and TypeScript SDK updates add beta capabilities for operating browsers and computers. Some existing Azure and Google Cloud customers receive Sonnet’s caching reduction over the following days.

The model also tightens cybersecurity restrictions compared with Haiku 4.5, including blocking penetration testing under its standard safeguards. Organizations seeking broader cybersecurity or biology access can apply through Anthropic’s verification programs.

For enterprise developers, the practical test is whether Haiku can complete a narrowly defined assignment reliably enough to justify repeated use. Its lower prices and reported gains make that test more attractive; its own benchmark results still support reserving more complex work for larger models.