Latest AI News

August 14, 2026 · Daily brief

OpenAI Ultrafast pushes GPT-5.6 Sol to 750 tokens/sec—on rented chips

Sovereignty angle
Frontier intelligence just got real-time—but only if you rent it. OpenAI doesn't own the chips, you don't own the API, and when Cerebras capacity fills up, your speed disappears. The faster the model, the tighter the leash.

OpenAI previewed Ultrafast, a Cerebras-powered tier pushing GPT-5.6 Sol to 14x normal speed and 750 tokens per second. The speed comes from a 750MW partnership—but it's invite-only API.

OpenAI and Cerebras unveiled Ultrafast mode on August 13, delivering GPT-5.6 Sol at speeds up to 750 output tokens per second—as much as 14 times faster than standard processing. The preview is invite-only through the OpenAI API, with no public pricing announced yet.

Speed through specialized silicon

The breakthrough comes from Cerebras' wafer-scale architecture, which consolidates massive compute, memory, and bandwidth onto a single giant chip. This eliminates the memory-to-processor bottlenecks that slow conventional GPU-based inference. The partnership, announced in January 2026, commits 750 megawatts of Cerebras capacity to OpenAI through 2028 in a deal reportedly valued at over $10 billion.

In testing, Ultrafast completed Humanity's Last Exam—2,500 graduate-level questions—in just over 11 hours compared to 78 hours for Claude Fable 5, with comparable accuracy. On GDP-Val knowledge-work tasks, it delivered a 5.6x end-to-end speedup. One OpenAI engineer said the speed feels like "genuinely cheating at my job," while another saw security investigations drop from hours to 10 minutes.

The infrastructure dependency

OpenAI positions Ultrafast for real-time workflows: voice, customer support, coding agents, financial research, and incident response. But access hinges entirely on Cerebras' rollout schedule and OpenAI's API terms. There's no path to self-host, no option to bring your own hardware, and pricing remains undisclosed.