OpenAI has switched on a new Ultrafast service tier for GPT-6.1 Sol, extending the speed-first processing mode it introduced for flagship GPT-6 Astra to the cheaper model that most developers actually route production traffic to. The tier is available today across the Responses API, Codex and ChatGPT Work, and it arrived on day four of the company's "28 days of updates" run, alongside near-instant steering for running Codex tasks and a Codex CLI 0.162.0 release.
Ultrafast is not a new model. It is a service_tier flag set on an existing gpt-6.1-sol call, and OpenAI's documentation is explicit about what it changes: the speed at which output tokens are generated, advertised as up to 8 times faster than the Sol Standard tier. The company describes the combination as "near-Astra intelligence at up to 8x faster speeds," and developer relations engineer Dominik Kundel framed the positioning as roughly Astra-level intelligence, eight times Sol's speed, at about 1.2 times Astra's cost.
The positioning is priced accordingly. Ultrafast costs $12 per million input tokens and $60 per million output tokens - six times Sol Standard's $2 and $10 - and sits above the mid-tier Fast option at $4 and $20. Inputs longer than 272,000 tokens rise to $24 and $90. In the API the tier is open to all developers under separate rate limits (1 million tokens per minute on the Build usage tier, 4 million on Launch and 40 million on Grow), and it supports both US and EU data residency.
Access narrows sharply outside the API. In ChatGPT Work and Codex, Ultrafast is limited at launch to the $500-per-month Pro plan, qualifying usage-based Enterprise accounts and credit-based Edu workspaces, with Enterprise administrators required to switch it on per user or workspace. Plus, the $100 and $200 Pro tiers and Business plans are excluded, and the consumer ChatGPT app is not part of the rollout at all.
The caveats matter. The 8x figure is a vendor claim: benchmark firm Artificial Analysis measures standard Sol at roughly 67 output tokens per second, below the category median, but no third party has yet published a tokens-per-second measurement for the Sol Ultrafast variant, and the "up to" is doing real work in OpenAI's sentence. The same applies to the original Astra Ultrafast tier, which the company demonstrated at roughly 300 tokens per second at its September developer day. It is also not OpenAI's first Ultrafast experiment: a preview built with Cerebras for GPT-5.6 Sol in August advertised 14x standard speed, and a later report that the mode actually runs on Nvidia silicon helped push Cerebras shares to a post-IPO low.
The economics are the real story. Artificial Analysis puts Sol's intelligence index at 52 with full reasoning, one point behind Astra, and OpenAI says Sol matches Astra on the long-horizon software engineering benchmark DeepSWE v1.1 at about a fifth of the cost, and trails it by 2.1 points on OSWorld 2.0 at roughly a seventh of the cost per task. On those numbers, Ultrafast is the cheapest way to buy raw speed in the lineup: spread across the officially advertised maximum, the premium tier delivers more tokens per dollar than either Standard or Fast, even at six times the sticker price.
OpenAI's guidance for actually getting that speed is blunt: use WebSockets, "especially for agentic applications that make many tool calls in quick succession," because a fresh HTTP connection per request can eat the latency gains entirely. For teams where response time is the binding constraint - live incident triage, agent navigation, interactive coding - the tier turns a throughput problem into a budget line. For everything else, Standard still does the same thinking at a sixth of the price.
Comments (0)
Log in to join the discussion
Log InNo comments yet