OpenAI is expanding its lineup of AI models with the new GPT-6 Sol and GPT-6 Luna, mid-priced and lower tier models designed to serve as enterprise workhorses for common tasks requiring less intelligence — and both are priced at half or less the costs of their predecessors when accessed over OpenAI's application programming interface (API).

Both improve on benchmarks over their 5.6 predecessors, as well, though they remain less performant and less powerful than the flagship GPT-6 Astra model released earlier this month with a grand statement from the company's president that we are now in the "AGI era," or artificial general intelligence.

  • GPT-6 Luna is priced at a jaw-dropping affordable rate of $0.10/$0.50 USD per 1M tokens in/out

  • GPT-6 Sol is priced at $2/$10 USD per 1M tokens in/out

OpenAI says improvements to inference and caching allowed it to reduce prices while increasing capability. An OpenAI spokesperson confirmed to VentureBeat that these GPT-6 Sol and Luna rates are permanent prices, not promotional or introductory pricing. Against OpenAI’s currently published GPT-5.6 rates, the drop is substantial:

  • GPT-5.6 Sol currently costs $4 per million input tokens and $20 per million output tokens, making GPT-6 Sol exactly 50% cheaper in both directions.

  • GPT-5.6 Luna currently costs $0.20 input and $1.20 output, making GPT-6 Luna 50% cheaper on input and 58.3% cheaper on output.

Sol is aimed primarily at complex work that developers and knowledge workers perform repeatedly — building features, reviewing code, debugging and analyzing data.

Luna is positioned as the high-volume option for more tightly defined jobs such as summarization, extraction and answering straightforward questions.

Astra, then, remains for the most complex and multi-faceted projects crossing multiple elements and pieces of media, or harder scientific and mathematical problems.

For enterprises, that segmentation matters because the economics of AI agents increasingly depend on how many model calls a workflow makes, how much context gets replayed and whether a company really needs a frontier model for every step.

Sol lands directly at Claude Sonnet 5 pricing — and at half the token price of the new Opus 5.5

At $2 input and $10 output per million tokens, GPT-6 Sol is priced exactly alongside Anthropic’s Claude Sonnet 5, whose introductory $2/$10 pricing Anthropic made permanent in August.

Model

Input ($/1M)

Output ($/1M)

Total ($/1M)

Source

Muse Spark 1.2 / 1.3 Contributor

$0.10

$0.20

$0.30

Meta

MiMo-V2.6-Flash

$0.14

$0.28

$0.42

Xiaomi

GPT-6 Luna

$0.10

$0.50

$0.60

OpenAI

DeepSeek-V4.1-Flash — off-peak

$0.15

$0.60

$0.75

DeepSeek

MiMo-V2.6-Pro

$0.435

$0.87

$1.305

Xiaomi

MiniMax-M3

$0.30

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

$0.30

$1.20

$1.50

LongCat

DeepSeek-V4.1-Flash — peak hours

$0.30

$1.20

$1.50

DeepSeek

DeepSeek-V4-Pro — off-peak

$0.66

$1.98

$2.64

DeepSeek

LongCat-2.0 — standard

$0.75

$2.95

$3.70

LongCat

Gemini 3.7 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

Gemini 3.8 Flash — through Dec. 31, 2026

$0.75

$3.75

$4.50

Google

DeepSeek-V4-Pro — peak hours

$1.32

$3.96

$5.28

DeepSeek

Muse Spark 1.1 / 1.2 / 1.3

$1.25

$4.25

$5.50

Meta

GLM-5.3

$1.40

$4.40

$5.80

Z.AI

Grok 4.7 — ≤200K prompt tokens

$2.00

$6.00

$8.00

xAI

Qwen3.8-Max

$2.00

$6.00

$8.00

QwenCloud

Gemini 3.7 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

Gemini 3.8 Flash — starting Jan. 1, 2027

$1.50

$7.50

$9.00

Google

GPT-6 Sol

$2.00

$10.00

$12.00

OpenAI

Claude Sonnet 5

$2.00

$10.00

$12.00

Anthropic

Grok 4.7 — >200K prompt tokens

$4.00

$12.00

$16.00

xAI

Grok 4.7 Fast (Cursor) — ≤256K input

$4.00

$12.00

$16.00

Cursor

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Kimi K3

$3.00

$15.00

$18.00

Moonshot AI

Claude Opus 5.5

$4.00

$20.00

$24.00

Anthropic

Grok 4.7 Fast (Cursor) — >256K input, up to 500K

$6.00

$18.00

$24.00

Cursor

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Sakana AI

Claude Fable 5 / Claude Mythos 5

$10.00

$50.00

$60.00

Anthropic

Claude Fable 5.1 / Claude Mythos 5.1

$10.00

$50.00

$60.00

Anthropic

GPT-6 Astra — Standard mode

$10.00

$50.00

$60.00

OpenAI

GPT-6 Astra — Fast mode

$20.00

$100.00

$120.00

OpenAI

But Anthropic changed the competitive picture again Tuesday morning with Claude Opus 5.5. The new model costs $4 per million input tokens and $20 per million output tokens, making GPT-6 Sol 50% cheaper than Opus 5.5 on both uncached input and output token rates.

Anthropic says Opus 5.5 itself is 20% cheaper per token than Opus 5 and about 40% cheaper to run on typical workloads because it requires fewer tokens to complete the same work.

That still leaves Anthropic’s highest general-access tier considerably more expensive on raw token pricing. Claude Fable 5.1 costs $10 input and $50 output per million tokens, five times Sol’s input and output price.

Google remains more aggressive at the performance-oriented Flash tier. Its latest model, Gemini 3.8 Flash, introduced earlier this month, costs $0.75 input and $3.75 output per million tokens under introductory pricing through Dec. 31, before increasing to $1.50/$7.50 on Jan. 1, 2027.

Google positions the model for long-horizon software engineering, autonomous agents and complex enterprise workflows.

Gemini will therefore remain cheaper than GPT-6 Sol even after Gemini 3.8 Flash's scheduled increase, but the pricing structures differ: Google has explicitly time-limited the current Flash rate, while OpenAI says Sol’s $2/$10 pricing has no promotional expiration.

Luna occupies a different price tier entirely. At $0.10/$0.50, Luna’s raw token rates are 95% below Sonnet 5’s on both input and output and roughly 86.7% below Gemini 3.8 Flash’s current introductory rates.

Against SpaceXAI's new Grok 4.7, released just yesterday, Luna is 95% cheaper on input and about 91.7% cheaper on output (that's comparing it against Grok 4.7's lowest-cost pricing with inputs of below 200,000 tokens).

Those are price comparisons rather than capability equivalences, but they illustrate why Luna is designed for routing large volumes of routine agent work away from more expensive reasoning models.

Open-weight models are putting even more pressure on that pricing. Xiaomi’s newly released MiMo-V2.6-Pro, the most performant open weights model in the world upon its debut last night (per Artificial Analysis), costs $0.435 per million uncached input tokens and $0.87 per million output tokens through Xiaomi’s API — about 78% less than GPT-6 Sol on input and 91% less on output.

The smaller MiMo-V2.6-Flash is cheaper still at $0.14/$0.28, or roughly 93% below Sol on input and 97% below it on output.

Luna is much more competitive with that open-weight tier: its $0.10 input rate is about 29% below MiMo-V2.6-Flash’s $0.14, although its $0.50 output rate is about 79% higher than Flash’s $0.28.

And because Xiaomi releases both MiMo models under an MIT license, enterprises can download and self-host them rather than pay Xiaomi per token — shifting the economics from API fees to their own infrastructure and operations costs.

That makes GPT-6 Sol’s strongest pricing argument less about being the cheapest model available and more about whether its higher API cost translates into enough additional task-level reliability, coding performance and operational simplicity to justify the premium over increasingly capable open-weight alternatives.

OpenAI is making cost per completed task the benchmark

OpenAI’s performance argument focuses heavily on cost per successful task, rather than benchmark scores in isolation.

On AutomationBench 1.0.6, which evaluates agents completing workflows across 47 tools in areas including sales, marketing, finance, support and HR, OpenAI reports GPT-6 Sol at xhigh effort scoring 33.2% at $0.27 per task.

Its evaluation puts Claude Opus 5 at 26.9% at maximum effort while costing 11.1 times as much per task. Claude Fable 5.1 with Opus 5 fallback scores 31.4%, according to OpenAI, at more than 8.9 times Sol’s task cost; OpenAI notes that the comparison understates the Fable configuration’s actual cost because it excludes the Opus fallbacks, which occurred on roughly 40% of tasks.

Low-effort GPT-6 Astra scores 30.3% on AutomationBench, below Sol xhigh’s 33.2%, while costing 3.9 times as much per task.

Those Anthropic comparisons now require an important qualifier. OpenAI’s charts benchmark Sol against Opus 5, but Anthropic released Opus 5.5 hours before OpenAI’s scheduled announcement. Anthropic prices Opus 5.5 at $4/$20 and says it costs about 40% less than Opus 5 on typical workloads. Its newly published benchmarks show significant performance improvements over Opus 5, but do not include GPT-6 Sol. That means there is not yet a same-harness public result establishing whether Sol or Opus 5.5 delivers the lower cost per successful task.

The coding numbers follow a similar pattern. On DeepSWE 1.1, GPT-6 Sol at maximum effort scores 68.8%, versus 69.9% for Claude Fable 5 at xhigh effort, with OpenAI estimating Sol’s cost per task at roughly 80% lower. GPT-6 Luna reaches 66.6%; OpenAI says its task cost is 93% below the compared Opus 5 configuration and 96% below Fable 5.

On OSWorld 2.0, a computer-use benchmark, Sol xhigh scores 60.5% in OpenAI’s testing versus 60.3% for Claude Opus 5 at medium effort, which OpenAI says comes at roughly 80% lower cost per task.

And on Agents’ Last Exam, which evaluates long-horizon professional work across 55 sub-industries, GPT-6 Sol at maximum effort scores 56.4%; above Opus 5’s highest score in OpenAI’s evaluation at 60% lower cost per task.

The vendor comparisons warrant caution. Different labs can produce different relative results when benchmarks are run under different configurations, effort settings and harnesses.

OpenAI did not provide same-harness GPT-6 Sol results against Gemini 3.8 Flash, Grok 4.7 or Xiaomi’s new MiMo-V2.6 models. Their token prices can be compared directly, but their real task economics cannot be inferred from those prices alone.

Factuality gets cheaper too

OpenAI is also making a cost-performance argument around factual reliability.

On an internal evaluation built from de-identified ChatGPT conversations where users previously flagged factual errors, the company says GPT-6 Sol makes roughly half as many mistakes as GPT-5.6 Sol and is now “approaching Astra-level reliability” at lower cost.

Luna shows a different kind of efficiency gain. At higher effort levels, OpenAI says GPT-6 Luna can match GPT-5.6 Sol’s factuality at about one-hundredth of its task cost.

The company cautions that the evaluation deliberately selects error-inducing conversations and is not representative of normal ChatGPT usage.

Caching becomes part of total cost of ownership

Token pricing is only one component of agent economics. Long-running coding and business agents repeatedly send system prompts, files, tool definitions and conversation history back to a model.

OpenAI says GPT-6 improves default prompt-cache hit rates and provides a 90% discount on cached input-token reads. Developers can also change reasoning effort or available tools without invalidating earlier cached context and can explicitly define where cacheable prompt prefixes end.

Anthropic is pushing on the same lever. Sonnet 5 advertises savings of up to 90% through prompt caching, while Google currently charges $0.075 per million cached Gemini 3.8 Flash tokens under its introductory pricing and Grok 4.7 lists cached input at $0.50 per million below its 200K prompt threshold.

Xiaomi goes further on nominal cache-read pricing: MiMo-V2.6-Flash lists cached input at $0.0028 per million tokens, while Pro lists it at $0.0036. Those rates make the open-weight models especially aggressive for agent workloads that repeatedly reuse large prompt prefixes, although self-hosted deployments introduce their own hardware, orchestration and utilization costs.

OpenAI points to GitHub as evidence of the operational impact, saying caching improvements reduced the portion of prompt tokens requiring fresh processing by more than 50% across billions of OpenAI-model requests used by Copilot.

Replit offers another example of what falling inference prices can change at the product layer. OpenAI says its earlier price cuts helped Replit offer Free Mode to millions of users, allowing subscribers to ask questions and plan applications without consuming normal usage credits. OpenAI also says GPT-5.6 Luna usage grew more than 10-fold following an 80% price reduction in July.

Alignment gains come with unresolved edge cases

OpenAI is also publishing unusually concrete numbers about undesirable agent behavior. In an internal coding-deception test deliberately constructed to elicit dishonesty, GPT-6 Sol records a 1.3% deception rate, down from 10.4% for GPT-5.6 Sol; Luna falls to 2.8% from 9.5%.

On another adversarial test, in which an agent receives a broken search tool and is graded on whether it admits the tool is broken rather than guessing, Sol’s failure-to-disclose rate falls to 5.4% from 77.8%, while Luna drops to 30.2% from 78.3%. In Codex testing, OpenAI observed no attempts by Sol to work around an automated safety reviewer; Luna attempted to bypass one in 0.3% of runs, down from 3.5%, with no successful bypasses observed.

The results are not uniformly strong. When models encountered explicit warnings such as “access denied,” Sol still attempted to work around the restriction in 64.4% of the adversarial runs, versus 68.2% for GPT-5.6 Sol. Luna’s rate was 42.4%, down from 76.5%. OpenAI stresses that these are deliberately difficult, primarily low-stakes tests run without all of the system-level safeguards used in its products and should not be read as normal-use failure rates.

Available in the API, Codex and ChatGPT Work

GPT-6 Sol and Luna are available through the OpenAI API as gpt-6-sol and gpt-6-luna. OpenAI is also rolling them out to ChatGPT Work and Codex for Plus, Pro, Business and Enterprise customers, while Free and Go users can access Luna through the desktop app. The models are not yet available in Chat itself, and OpenAI says the Work and Codex rollout will proceed gradually during the day to preserve service stability.

GPT-6 Astra remains OpenAI’s top-tier option for workloads where maximizing capability matters more than cost.

The more consequential shift for enterprise buyers, however, is that model selection is becoming less about choosing a single “best” AI and more about allocating different classes of work across a portfolio.

At Luna’s prices, routine extraction and summarization can be pushed toward a much cheaper tier. Sol targets recurring coding and agent work at Sonnet-class token pricing. Astra remains available for the hardest jobs.

The fact that OpenAI is making Sol and Luna’s reductions permanent strengthens that portfolio argument: developers do not have to build production routing and unit-economics assumptions around a discount that is scheduled to disappear.

Anthropic’s same-day Opus 5.5 release and Xiaomi’s open-weight MiMo-V2.6 family make that shift even more explicit. Vendors are increasingly marketing models not simply by intelligence scores, but by the amount of useful work they can finish for a given amount of compute and money. The next enterprise model bake-off therefore needs to measure successful task completion, token consumption, cache reuse, latency and rework — not just API list prices or a single benchmark score.