OpenAI and Cerebras unveiled Ultrafast mode on August 13, delivering GPT-5.6 Sol at speeds up to 750 output tokens per second—as much as 14 times faster than standard processing. The preview is invite-only through the OpenAI API, with no public pricing announced yet.
Speed through specialized silicon
The breakthrough comes from Cerebras' wafer-scale architecture, which consolidates massive compute, memory, and bandwidth onto a single giant chip. This eliminates the memory-to-processor bottlenecks that slow conventional GPU-based inference. The partnership, announced in January 2026, commits 750 megawatts of Cerebras capacity to OpenAI through 2028 in a deal reportedly valued at over $10 billion.
In testing, Ultrafast completed Humanity's Last Exam—2,500 graduate-level questions—in just over 11 hours compared to 78 hours for Claude Fable 5, with comparable accuracy. On GDP-Val knowledge-work tasks, it delivered a 5.6x end-to-end speedup. One OpenAI engineer said the speed feels like "genuinely cheating at my job," while another saw security investigations drop from hours to 10 minutes.
The infrastructure dependency
OpenAI positions Ultrafast for real-time workflows: voice, customer support, coding agents, financial research, and incident response. But access hinges entirely on Cerebras' rollout schedule and OpenAI's API terms. There's no path to self-host, no option to bring your own hardware, and pricing remains undisclosed.