OpenAI's New Ultrafast Mode Runs Its Best Model 14 Times Faster
A new OpenAI API tier powered by Cerebras runs GPT-5.6 Sol at up to 750 output tokens per second. Same model, same quality, roughly fourteen times the speed, for a small group of customers first.
OpenAI previewed a new service tier on Thursday called Ultrafast, which runs GPT-5.6 Sol, its most capable model, at up to fourteen times the usual speed. In concrete numbers, that is up to 750 output tokens per second. Tokens are the chunks of text a model produces, roughly three quarters of a word each, so 750 per second is faster than any human can read. The hardware behind it comes from Cerebras, the chip company whose processors are built on a single enormous wafer instead of many small chips.
The point OpenAI is making is that this is the same model, not a smaller one. Until now, the way to get a fast answer was to reach for a lighter model and accept that it would be worse at hard problems. Ultrafast keeps the intelligence and removes the wait. OpenAI reports a 5.6x end-to-end speedup on GDPval, a benchmark built around economically valuable knowledge work, with no measurable drop in quality. Cerebras claims the setup is five times faster than Claude Opus 4.8 and eleven times faster than Claude Fable 5.
The early customers describe what changes when the delay disappears. Podium uses it in a voice product, where waiting three seconds for a reply on a phone call is the difference between a conversation and an awkward silence. The financial research firm Rogo says complex research starts to feel like a real-time interaction. Inside OpenAI, engineers use it during outages: when an alert fires, the model reads logs, analyses traces, and suggests next checks while the system is still misbehaving, rather than after. Researchers describe a batch of overnight experiments collapsing into several rounds during a single working day.
What is actually going on here
Speed is quietly becoming its own axis of competition. For two years the industry raced on capability, and the models got smart enough that for many jobs the remaining bottleneck is not “can it do this” but “will it answer before the moment passes”. A model that takes forty seconds cannot hold a phone conversation, cannot help during an outage, cannot sit inside a checkout flow. Getting there needs specialised silicon, which is why OpenAI is leaning on Cerebras rather than doing it in-house, and it is a reminder that the AI story is as much about chips as about models. A fair caveat: this is a limited preview for selected customers, OpenAI has not published pricing, and “up to 750 tokens per second” is a ceiling, not an average.
What this means for you: nothing today, and that is fine. This is an API tier for businesses, not a button in ChatGPT. What it does tell you is where the products you use are heading. Voice assistants that stop leaving gaps, support chats that resolve something instead of stalling, coding tools that keep pace with your typing: those are all speed problems, and they are being solved. When your bank’s chatbot suddenly stops feeling like a hold queue, this is the kind of change underneath it.
Sources
Suno Studio 2.0 Lets You Talk to a Music Studio Like It Is a Bandmate
Suno turned its music generator into a chat-driven production tool: build instruments and plugins by typing, import MIDI, export unlimited 32-bit multitrack. The export policy sits oddly next to last week's anti-spam limits.