OpenAI Launches Ultrafast Mode - GPT-5.6 Sol at 14x Speed via Cerebras in Limited Preview

Powered by Cerebras, Ultrafast runs OpenAI's most intelligent model at 750 output tokens per second - targeting incident response, voice, and financial research where latency determines usability.

Saganote
Saganote ·
3 Min Read

14 times faster. OpenAI Ultrafast launched August 13 as a new API service tier running GPT-5.6 Sol at up to 750 output tokens per second - over an order of magnitude ahead of standard inference speeds and powered by a chip partnership with Cerebras. Access is in limited preview today for a select group of API customers while OpenAI scales capacity.

Cerebras Hardware Runs GPT-5.6 Sol at 750 Tokens per Second - 14x the Standard Rate

Standard GPT-5.6 Sol is already fast. OpenAI Ultrafast runs the same model on Cerebras hardware at up to 14x that rate, reaching 750 output tokens per second - a throughput class that changes what frontier intelligence can do in workflows where latency determines usability. GPT-5.6 Sol launched in July as OpenAI's flagship intelligence tier across three variants - Sol, Terra, and Luna; Ultrafast does not modify the model, only its inference substrate. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," OpenAI wrote at launch - Ultrafast is the argument that the tradeoff is no longer necessary.

Incident Response, Voice Support, and Financial Research Are the First Target Workflows

Five categories, one common requirement: answers before conditions change. OpenAI named incident response, financial market analysis, customer support and voice, e-commerce personalization, and live research experimentation as the primary Ultrafast use cases in its launch post. Podium was among the first external testers. "The speed completely changes the call experience for the more complex work," said Courtland Lykins, Podium's Product Lead for Voice AI. GPT-Live, OpenAI's full-duplex voice model built for real-time conversations, targets the same latency-sensitive territory; Ultrafast gives Sol the throughput to compete there without dropping to a lighter model.

Research iteration speed was the internal case. OpenAI's own engineering teams used Ultrafast to compress overnight experiment batches into same-day loops - launching, reviewing results, adjusting, and running again within a single workday instead of across days. Alex Wang, Applied AI at Rogo, which used Ultrafast for financial research, summarized the practical shift: "Speed doesn't just make the product feel better. It changes what people can realistically use it for."

Limited Preview Now - No Pricing or General Availability Date Announced

No pricing and no timeline for general availability. OpenAI Ultrafast is live today for a handpicked group of API customers; the company opened a sign-up form for teams wanting notification as access expands alongside Cerebras capacity. OpenAI cut GPT-5.6 Sol's thinking budget 87% within days of the model's launch to manage compute costs; Ultrafast moves in the opposite direction, adding Cerebras hardware as a separate inference path for customers who need the speed and can pay for it. Jane Street, Podium, Basis, and Rogo appear among the early access group; OpenAI has not named a date for broader rollout.


Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, cybersecurity, startups, software, consumer technology, and innovation. Our editorial team researches, writes, and reviews original news, analysis, and explainers to provide accurate, timely, and well-sourced coverage of the technology industry.