Anthropic today released Claude Opus 5.5, its new frontier model aimed at long-running coding agents, research, and professional knowledge work and it both outperforms and undercuts the cost (over application programming interface, or API) of Anthropic's prior flagship models Mythos 5.1 and Fable 5.1, released just a few weeks earlier.
Anthropic prices Opus 5.5 over the API at $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5, while claiming typical workloads cost about 40% less overall because the model also uses fewer tokens to finish tasks.
Anthropic also cut cache-write pricing by 20%, from $6.25 to $5 per million tokens, while reducing cache-read costs even more sharply, from $0.50 to $0.20 per million tokens.
Impressively, Opus 5.5 does this while also achieving some of the highest benchmarks yet from any Anthropic model, including on agentic coding (Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0) knowledge work (GDPval-AA v2.1), scientific and multidisciplinary reasoning (Terminal-Bench-Science 0.1³, Humanity's Last Exam), and more.
Opus 5.5's benchmarks largely exceed Fable 5.1, Anthropic's prior flagship for the general public, even though Fable 5.1 is 150% more expensive over API.

The release arrives as frontier model vendors increasingly compete on the amount of useful work completed per dollar rather than benchmark scores alone.
Model | Input ($/1M) | Output ($/1M) | Total ($/1M) | Source |
Muse Spark 1.2 / 1.3 Contributor | $0.10 | $0.20 | $0.30 | |
MiMo-V2.6-Flash | $0.14 | $0.28 | $0.42 | |
GPT-6 Luna | $0.10 | $0.50 | $0.60 | |
DeepSeek-V4.1-Flash — off-peak | $0.15 | $0.60 | $0.75 | |
MiMo-V2.6-Pro | $0.435 | $0.87 | $1.305 | |
MiniMax-M3 | $0.30 | $1.20 | $1.50 | |
LongCat-2.0 — limited-time promo | $0.30 | $1.20 | $1.50 | |
DeepSeek-V4.1-Flash — peak hours | $0.30 | $1.20 | $1.50 | |
DeepSeek-V4-Pro — off-peak | $0.66 | $1.98 | $2.64 | |
LongCat-2.0 — standard | $0.75 | $2.95 | $3.70 | |
Gemini 3.7 Flash — through Dec. 31, 2026 | $0.75 | $3.75 | $4.50 | |
Gemini 3.8 Flash — through Dec. 31, 2026 | $0.75 | $3.75 | $4.50 | |
DeepSeek-V4-Pro — peak hours | $1.32 | $3.96 | $5.28 | |
Muse Spark 1.1 / 1.2 / 1.3 | $1.25 | $4.25 | $5.50 | |
GLM-5.3 | $1.40 | $4.40 | $5.80 | |
Grok 4.7 — ≤200K prompt tokens | $2.00 | $6.00 | $8.00 | |
Qwen3.8-Max | $2.00 | $6.00 | $8.00 | |
Gemini 3.7 Flash — starting Jan. 1, 2027 | $1.50 | $7.50 | $9.00 | |
Gemini 3.8 Flash — starting Jan. 1, 2027 | $1.50 | $7.50 | $9.00 | |
GPT-6 Sol | $2.00 | $10.00 | $12.00 | |
Claude Sonnet 5 | $2.00 | $10.00 | $12.00 | |
Grok 4.7 — >200K prompt tokens | $4.00 | $12.00 | $16.00 | |
Grok 4.7 Fast (Cursor) — ≤256K input | $4.00 | $12.00 | $16.00 | |
GPT-5.4 | $2.50 | $15.00 | $17.50 | |
Kimi K3 | $3.00 | $15.00 | $18.00 | |
Claude Opus 5.5 | $4.00 | $20.00 | $24.00 | |
Grok 4.7 Fast (Cursor) — >256K input, up to 500K | $6.00 | $18.00 | $24.00 | |
Sakana Fugu Ultra (≤272K) | $5.00 | $30.00 | $35.00 | |
Claude Opus 5.5 — Fast mode | $8.00 | $40.00 | $48.00 | |
Claude Fable 5.1 / Claude Mythos 5.1 | $10.00 | $50.00 | $60.00 | |
GPT-6 Astra — Standard mode | $10.00 | $50.00 | $60.00 | |
GPT-6 Astra — Fast mode | $20.00 | $100.00 | $120.00 |
Also today, OpenAI released GPT-6 Sol and GPT-6 Luna, two lower-cost members of its GPT-6 family. Sol costs $2 per million input tokens and $10 per million output tokens — exactly half Opus 5.5’s base rates.
Luna drops to just $0.10/$0.50, or 97.5% below Opus 5.5 on both raw input and output pricing.
The same-day launches sharpen an increasingly important enterprise question: not which vendor has the single most capable model, but how much useful autonomous work each model can complete for a given amount of money.
Agentic coding is the central Opus 5.5 pitch
Anthropic’s strongest performance claims center on software engineering tasks that require an agent to work across a codebase for extended periods rather than answer isolated coding questions.
The company reports 66.4% on Terminal-Bench 4.0, compared with 55.8% for Fable 5.1 and 52.3% for Opus 5. On FrontierCode v1.1 Main, Anthropic reports 54.4% for Opus 5.5, versus 50.3% for Fable 5.1 and 48.0% for Opus 5. And on CursorBench 4.0, Opus 5.5 reaches 57.8%, compared with 51.8% and 46.6%, respectively.
Impressively, Opus 5.5 does this while also achieving some of the highest benchmark scores Anthropic has reported, including in agentic coding, knowledge work, and scientific and multidisciplinary reasoning. Its scores largely exceed Fable 5.1, Anthropic’s prior flagship for general users, while Opus 5.5 is 60% cheaper on base API token pricing.
Anthropic also points to large real-world coding workloads. One early tester completed a 680,000-line code migration in less than a day. Another used Opus 5.5 to audit and repair a 200,000-line codebase in under three hours, where Opus 5 reportedly took more than 20 hours and consumed 2.5 times as many tokens.
In an internal experiment translating HAProxy from C to Rust, Anthropic says both Opus 5.5 and Fable 5.1 passed nearly all of the project’s regression tests, but Opus 5.5 finished in 9.5 hours versus 12 hours and cost 51% less.
GitHub Chief Product Officer Mario Rodriguez said in a statement provided by Anthropic that Opus 5.5 used “among the fewest tokens and steps” GitHub measured across Copilot CLI and VS Code testing, while solving more terminal tasks than Opus 5 in fewer than half the steps. Other early testers including Lovable, Spotify and Optiver similarly reported reductions in turns, output tokens or execution time.
Those are vendor-selected customer results rather than independent evaluations, but they point to the operational metric Anthropic wants enterprise buyers to watch: how much orchestration, retrying and token consumption occurs before the work is actually finished.
GPT-6 Sol changes the cost comparison
OpenAI is making essentially the same argument with GPT-6 Sol and Luna — but from a substantially lower raw API price.
At $2/$10 per million input/output tokens, Sol is positioned for repeated complex work including coding, debugging, feature development and data analysis.
Luna, at $0.10/$0.50, targets high-volume tasks such as extraction, summarization and straightforward question answering. OpenAI’s top-tier GPT-6 Astra remains reserved for workloads where maximum capability matters more than cost.
That makes Sol the more relevant price competitor for Opus 5.5: its uncached token prices are exactly 50% lower.
Caching narrows part of that gap. OpenAI says GPT-6 provides a 90% discount on cached input reads, which implies $0.20 per million cached Sol tokens — the same nominal cache-read rate Anthropic lists for Opus 5.5. Luna’s corresponding rate is just $0.01 per million.
OpenAI also says GPT-6 can preserve cached context when developers change reasoning effort or enable and disable tools, potentially increasing cache reuse in long-running agents.
Still, raw pricing does not establish total workload cost. A cheaper model can become more expensive if it needs more calls, produces more tokens or fails often enough to require retries.
And right now, there is no clean public same-harness comparison between GPT-6 Sol and Claude Opus 5.5 across the major agentic benchmarks.
OpenAI reports GPT-6 Sol at 33.2% on AutomationBench 1.0.6 at xhigh effort for $0.27 per task. Anthropic, meanwhile, reports an early-access Zapier evaluation of Opus 5.5 at 40.0%.
On their face, the figures favor Opus 5.5 on task completion, while Sol carries the dramatically lower price.
But OpenAI notes that competitor results in its release were taken from public reports rather than generated in the same evaluation environment, so the two numbers should not be treated as a controlled head-to-head result.
OpenAI also says Sol can match Claude Fable 5.1 at xhigh effort on FrontierCode 1.1 Main at much lower cost. Anthropic’s own Opus 5.5 release places the new model above Fable 5.1 on FrontierCode, but again, the two companies have not published a common Opus 5.5-versus-Sol run from the same harness and settings.
Luna is less of a direct Opus rival and more of a potential routing alternative. OpenAI reports 66.6% on DeepSWE v1.1 at maximum effort and says it reaches performance comparable to older high-end Claude configurations at a fraction of their task cost.
That creates another option for enterprises that do not need to send every step in an agent workflow to a frontier-priced model.
Knowledge work follows the same efficiency strategy
Anthropic is also positioning Opus 5.5 for financial analysis, research, document production and other professional workflows.
On GDPval-AA v2.1, a benchmark of work across 44 occupations, Anthropic reports 1846 Elo for Opus 5.5, compared with 1735 for Fable 5.1 and 1708 for Opus 5.
In an internal research evaluation, an automated grader rejected any report containing an invented figure or quotation. Anthropic says 16 of 18 Opus 5.5 reports cleared its quality threshold across different effort settings, while neither Fable 5.1 nor Opus 5 cleared the threshold in any attempt. Because this is an Anthropic-designed internal test, it should be read as supporting evidence rather than an independent benchmark.
That distinction matters as vendors emphasize task economics. GPT-6 Sol’s launch similarly focuses on professional workflows and cost per successful task rather than simply publishing isolated intelligence scores. OpenAI says Sol at xhigh effort can even outperform low-effort GPT-6 Astra on AutomationBench while costing substantially less per run.
The emerging pattern is model routing rather than one-model standardization: expensive models for the hardest ambiguous work, lower-priced models for repeated execution, and very inexpensive tiers such as Luna for tightly scoped sub-tasks.
Developers get lower Opus prices, but also breaking changes
Opus 5.5 has a 1 million-token context window, supports up to 128,000 output tokens synchronously and uses the API model ID claude-opus-5-5. Anthropic also supports outputs of up to 300,000 tokens through its Message Batches API under a beta configuration.
But Opus 5.5 is not a drop-in replacement for every Opus 5 application.
Adaptive thinking is now always enabled, with developers controlling depth through the effort parameter rather than switching reasoning off. Forced tool use can return an error; thinking blocks are tied to the model and conversation that produced them; the older computer_20251124 computer-use tool is rejected on the Claude API and Google Cloud; and applications that expose text between tool calls as progress updates may need to change how they process thinking blocks.
For production agent systems, those are meaningful changes affecting tool routers, persisted conversation state, regression tests and user interfaces.
Safety controls can also alter runtime behavior
Anthropic is coupling Opus 5.5’s capability gains with additional controls around cybersecurity, biology and model distillation.
Some cybersecurity requests can be transparently routed to Opus 4.8, while safeguards for sensitive biology work can similarly result in fallback behavior. Anthropic says vetted organizations will be able to obtain broader access through its verification programs.
Anthropic reports that Opus 5.5 performs better than recent Claude models across nearly every measure in its roughly 2,000-scenario automated behavioral audit.
In a new containment evaluation, it says Opus 5.5 attempted to circumvent boundaries roughly 85% less often than Opus 5 or Claude Mythos 5.1. The company also cautions that reliably detecting every failure before deployment remains an unsolved problem.
OpenAI is likewise emphasizing alignment improvements in Sol and Luna, reporting lower rates of misleading claims about completed coding work and other undesirable agent behaviors. It stresses that these evaluations are deliberately adversarial and should not be interpreted as normal-use failure rates.
The competition is moving toward cost per completed task
Claude Opus 5.5 is available through Anthropic’s API and across AWS, Google Cloud and Microsoft’s ecosystem.
GPT-6 Sol and Luna are available through OpenAI’s API as gpt-6-sol and gpt-6-luna, and are rolling out through ChatGPT Work and Codex.
The competitive picture after Tuesday’s dual launches is therefore more complicated than a benchmark leaderboard.
Anthropic has pushed high-end Claude capability down dramatically in price: Opus 5.5 beats its previous general-access flagship on most of the benchmarks Anthropic reports while costing 60% less per base API token. OpenAI has responded at another point on the curve, with Sol charging half of Opus 5.5’s uncached token price and Luna establishing an even cheaper tier for work that does not need flagship-level reasoning.
Until both new models appear in controlled, same-harness evaluations, it is premature to reduce the choice to a single price-performance winner.
For enterprise teams, the more useful bake-off is increasingly internal: successful task completion, token consumption, cache reuse, latency, rework and human intervention across the workloads they actually plan to deploy.
The model war is no longer only about who can produce the highest score. It is increasingly about who can turn frontier intelligence into the most completed work per dollar.
