OpenAI previews Ultrafast, an API tier powered by Cerebras that runs GPT-5.6 Sol up to 14× faster and generates up to 750 output tokens per second
OpenAI is previewing a new way to run its most capable GPT-5.6 model at dramatically higher speeds. The company says its new Ultrafast …
9to5MacZac Hall
Context & Ripple Effects
OpenAI has been building a performance ladder around serving speed: its GPT-5.3-Codex-Spark research preview emphasized much faster code generation, while GPT-5.5 was positioned as maintaining GPT-5.4-level per-token latency at a higher capability level. Ultrafast extends that focus to GPT-5.6 Sol through an API tier backed by Cerebras rather than a smaller model variant.
First-order effects
API customers can access GPT-5.6 Sol through a tier OpenAI says delivers up to 14× faster performance and up to 750 output tokens per second.
Cerebras becomes the named infrastructure partner behind a customer-facing OpenAI API offering, tying its serving technology directly to OpenAI's high-speed tier.
Second-order effects
OpenAI's faster GPT-5.6 Sol endpoint raises the serving-speed benchmark for API rivals, particularly for coding and agent-style workloads that were already targeted by GPT-5.4 mini and nano.
OpenAI can differentiate API access by runtime speed as well as model capability, making infrastructure partners such as Cerebras more consequential to product positioning.
Third-order effects
The pattern points toward inference tiers becoming a distinct product layer: the same model family can be sold on different latency and throughput characteristics, not only on intelligence or size.
If OpenAI continues pairing model releases with specialized serving paths, competition will increasingly center on the integrated model-and-inference stack rather than model quality alone.
The trend: Frontier-model providers are turning low-latency inference into a differentiated API product, linking model roadmaps more tightly to specialized compute partners.
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows. [video]
Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras. GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model — up to 14× faster than the same model on Standard processing. It speedran Humanity's Last [video…
Prediction: we will soon enter an era where the latency of agentic work is bottlenecked not by LLM inference speed but by tool use. Shell commands, build pipelines, computer use, web data extraction, and everything else that happens in the background is way slower than it needs …
We ran GPT-5.6 Sol on Ultrafast mode through Humanity's Last Exam. 2,500 questions across chemistry, economics, literature - questions typically only PhDs could answer. It finished the entire benchmark in 11 hours and 11 minutes. Nearly 7× faster than Claude Fable 5. Frontier
GPT-5.6 Sol Ultrafast is here!! 750 tokens per second honestly feels instantaneous for most workloads. The bottleneck now moves to the tools themselves.
Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras. We gave @OpenAI's GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts. Ultrafast: 1 min 50 seconds Standard: 12 min 20 seconds Same result, [v…
14x the speed. @Cerebras and @OpenAI. Next week is Cerebras' event in San Francisco. My wife is helping put that on, so have an inside scoop and will be there. Marriage with benefits. :-)
Frontier intelligence meets Cerebras speed. — OpenAI just announced Ultrafast, a new service tier running GPT-5.6 Sol on Cerebras at up to 750 output tokens per second. …