OpenAI unveiled GPT-6.1 Sol, an upgraded version of GPT-6 Sol that OpenAI says approaches its flagship GPT-6 Astra across coding, computer use and professional workflows while charging one-fifth of Astra’s standard input and output token prices.
Alongside it, OpenAI is formally introducing Ultrafast, a premium inference tier offering up to 8X faster token generation in Codex and 6X in the API, reaching as much as 300 tokens per second.
At that speed, OpenAI’s Ultrafast tier would sit firmly in the high-speed end of today’s model market, but it would not be the outright throughput leader.
Independent, third-party benchmarking firm Artificial Analysis currently measures Google’s Gemini 3.5 Flash at about 201 tokens/sec, while specialized speed-focused models go much higher: Mercury 2 at roughly 769 tokens/sec and Celeris-1 at about 1,491 tokens/sec.

The distinction is that OpenAI is offering that 300-token/sec ceiling on its frontier GPT-6-class models, whereas the absolute speed leaders tend to be models optimized specifically for ultra-high-throughput inference.
The speed comes with a substantial premium: API use costs 6X the corresponding model’s standard rate. Ultrafast is available immediately for GPT-6 Astra, while the company says a version for 6.1 Sol is coming soon.
Together, the releases create a wider price-performance spectrum for enterprise developers. Organizations can run increasingly sophisticated autonomous work on a comparatively inexpensive Sol model, or pay considerably more when response latency rather than token cost becomes the limiting factor.
GPT-6.1 Sol keeps Sol pricing while moving toward Astra performance
The most important pricing detail may be what OpenAI didn’t raise.
GPT-6.1 Sol costs $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens.
Its uncached input and output rates are the same as GPT-6 Sol, while cached input falls from $0.20 to $0.10 per million tokens, a 50% reduction. OpenAI’s current developer documentation lists GPT-6 Sol at $2 input, $0.20 cached input and $10 output.
That makes the generational upgrade unusually straightforward from a budgeting perspective: developers can move from GPT-6 Sol to GPT-6.1 Sol without accepting a higher basic token rate.
The comparison with Astra is more dramatic. OpenAI currently charges $10 per million input tokens, $1 per million cached input tokens and $50 per million output tokens for GPT-6 Astra. GPT-6.1 Sol therefore costs exactly one-fifth as much for standard uncached input and output, while its $0.10 cached-input rate is one-tenth Astra’s.
For context, OpenAI’s currently discounted GPT-5.6 Sol costs $4 per million input tokens and $20 per million output tokens, meaning GPT-6.1 Sol is also half the present token price of that older Sol-generation model. OpenAI’s pricing documentation says GPT-5.6 Sol promotional pricing remains available at least through November 21, 2026.
Model / tier | Input / 1M | Cached input / 1M | Output / 1M |
GPT-5.6 Sol, current promotional rate | $4 | $0.40 | $20 |
GPT-6 Sol | $2 | $0.20 | $10 |
GPT-6.1 Sol | $2 | $0.10 | $10 |
GPT-6 Astra | $10 | $1 | $50 |
That cached-input reduction is particularly relevant for enterprise agents. Long-running coding agents, research systems and business-process automations frequently reuse system instructions, repository context, policies, reference documents and other prompt material between calls. Cutting the read cost of that reusable context can matter more than headline input pricing for workloads that execute hundreds or thousands of model turns.
OpenAI says Sol is getting much closer to Astra's performance
OpenAI backs its pricing argument with a series of its own evaluations showing GPT-6.1 Sol narrowing the gap with Astra. On DeepSWE v1.1, which evaluates long-running software-engineering work in real codebases, OpenAI says GPT-6.1 Sol matches Astra at roughly one-fifth the cost and beats GPT-6 Sol’s best result by 6.4 percentage points.
On GDP.pdf, a benchmark involving complicated professional documents containing charts, tables, diagrams and fine-print information, OpenAI reports that GPT-6.1 Sol scores above Anthropic’s Claude Opus 5.5 with fallbacks while costing less than half as much per task. It also approaches Astra at roughly one-fifth its cost.
The results may be more directly relevant to enterprise agent deployments on AutomationBench, which tests end-to-end tasks spanning 47 tools across sales, marketing, operations, support, finance and HR. OpenAI says GPT-6.1 Sol beats Opus 5.5 by 2.2 percentage points at medium reasoning effort while costing roughly one-third as much, and improves 4.8 percentage points over GPT-6 Sol at the same setting.
Computer use follows the same pattern. On the OSWorld 2.0 offline set, GPT-6.1 Sol comes within 2.1 percentage points of Astra at maximum reasoning effort, according to OpenAI, while costing roughly one-seventh as much per task.
Those are OpenAI-run evaluations rather than independent production tests. The company notes that its evaluations are conducted in its research environment or API and can differ somewhat from production ChatGPT. But the positioning is clear: Astra remains the maximum-capability option, while Sol is becoming the model OpenAI wants developers to use when the economics of running agents repeatedly matter as much as absolute benchmark performance.
Ultrafast makes inference speed a premium product
The second half of OpenAI’s announcement attacks a different bottleneck. Standard frontier models can be capable enough to complete a workflow while still being too slow for applications where a human is waiting for every response. OpenAI’s answer is Ultrafast, essentially a premium inference lane.
OpenAI says the tier reaches up to 300 generated tokens per second, with up to 8X faster generation in Codex and up to 6X faster generation through the API. API usage costs 6X the standard pricing.
That is considerably more aggressive than OpenAI’s existing Fast mode, formerly called Priority processing. OpenAI’s current API pricing table charges twice standard rates for GPT-6 family models: for example, GPT-6 Sol rises from $2/$10 input/output under Standard processing to $4/$20 in Fast mode, while Astra rises from $10/$50 to $20/$100.
Model / service tier | Speed | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens | Price vs. Standard |
GPT-6 Sol — Standard | Standard | $2 | $0.20 | $10 | 1X |
GPT-6 Sol — Fast mode | Faster than Standard | $4 | $0.40 | $20 | 2X |
GPT-6 Astra — Standard | Standard | $10 | $1 | $50 | 1X |
GPT-6 Astra — Fast mode | Faster than Standard | $20 | $2 | $100 | 2X |
GPT-6 Astra — Ultrafast | Up to 6X faster in API; up to 8X in Codex; up to 300 tokens/sec | $60* | $6* | $300* | 6X |
GPT-6.1 Sol — Standard | Standard | $2 | $0.10 | $10 | 1X |
GPT-6.1 Sol — Ultrafast | Up to 6X faster in API; up to 8X in Codex; up to 300 tokens/sec | $12* | $0.60* | $60* | 6X |
* Derived pricing: OpenAI says Ultrafast API usage costs 6X standard pricing, but its DevDay materials do not separately publish every Ultrafast line-item rate. The Astra and GPT-6.1 Sol Ultrafast figures above are therefore calculated from OpenAI’s published standard prices and announced 6X multiplier. OpenAI says Ultrafast costs three times as much as Fast mode and six times as much as Standard for the same base model.
Ultrafast therefore costs three times as much as Fast mode and six times as much as Standard processing, assuming the same base model.
For GPT-6 Astra, the announced 6X multiplier works out to $60 per million input tokens and $300 per million output tokensversus $10/$50 for Standard and $20/$100 for Fast mode.
For GPT-6.1 Sol, once Ultrafast arrives, the same announced multiplier would imply $12 input and $60 output per million tokens, with $0.60 cached input if the multiplier applies uniformly to its published $0.10 cached-input rate.
OpenAI’s DevDay materials state the 6X multiplier rather than separately listing those GPT-6.1 Sol Ultrafast token rates, so those figures are derived from the announced pricing rather than independently published line-item prices.
The economics mean Ultrafast is unlikely to replace Standard inference for background agents, document processing or other throughput-oriented jobs. Its obvious targets are applications where waiting carries its own cost: interactive coding, customer support, financial analysis, incident response and other human-in-the-loop systems.
OpenAI has already experimented with this idea. In August, it previewed GPT-5.6 Sol Ultrafast on Cerebras hardware, reaching as much as 750 output tokens per second and up to 14X Standard speed for a limited group of customers. OpenAI highlighted financial research, commerce, support and incident response as early use cases.
The DevDay rollout turns that experiment into a broader product tier, though the newer GPT-6 models trade some of the preview’s maximum headline speed for access to newer intelligence.
Availability — and a much larger DevDay push
GPT-6.1 Sol is available through the API as gpt-6.1-sol and for Plus, Pro, Business, Enterprise and Edu customers in ChatGPT Work and Codex. OpenAI explicitly says it is not yet available in regular Chat.
GPT-6 Astra Ultrafast is available now through the API and in ChatGPT Work and Codex for Pro 500 and Enterprise customers. GPT-6.1 Sol Ultrafast is due in the coming days.
The API reference also exposes an access-controlled service_tier=ultrafast option for supported models.
The model releases form only one part of a much larger DevDay push. OpenAI also announced its Dots always-on agents; Private Intelligence privacy infrastructure; computer use in the Agents API; AWS Bedrock Managed Agents powered by OpenAI; new Codex cloud, CLI, code-review and security capabilities; the Decisions API; expanded plugins and MCP event-triggered automations; collaborative ChatGPT Spaces, Pages and slides; team tasks; ChatGPT integrations for Slack and Microsoft Teams; a Meetings plugin; Sign in with ChatGPT; a new Pro 500 subscription; and an OpenAI Marketplace for enterprise customers.
For enterprise technical leaders, though, GPT-6.1 Sol and Ultrafast expose perhaps the more consequential shift underneath those individual products. OpenAI is increasingly asking developers to optimize not simply for which model is smartest, but across three separate variables: intelligence, cost and latency.
GPT-6.1 Sol pushes capable agent execution sharply down the cost curve. Ultrafast moves in the opposite direction, asking customers to pay a large premium when every second matters.
That gives enterprise architects a more granular decision to make — and potentially a more economical one: use expensive inference only where latency creates business value, while moving the much larger volume of autonomous work onto models whose capability is increasingly close to the frontier without carrying frontier prices.
