OpenAI previews 'Ultrafast' GPT-5.6 — 14x the speed, powered by Cerebras chips instead of GPUs
OpenAI is previewing an Ultrafast mode for GPT-5.6 Sol that runs up to 14 times faster than standard processing — around 750 output tokens a second — by routing inference through Cerebras hardware instead of GPUs, with no price or general release date confirmed yet.
23 August 2026
OpenAI began previewing “Ultrafast” mode for GPT-5.6 Sol in mid-August — a tier that runs the model up to 14 times faster than standard processing, generating roughly 750 output tokens a second. The speed comes from Cerebras’ wafer-scale chip hardware rather than the GPUs the rest of the industry runs on, under a partnership the two companies have been building out through the year. It’s a limited preview available to select customers through the API only, with teams testing it on coding, customer support, commerce and financial-research workloads — and no price or wider release date attached yet.
The number that matters here isn’t the headline multiplier, it’s what it’s competing on. Most of 2026’s AI-model news cycle has been a price war — Gemini 3.7 Flash’s cuts, DeepSeek V4 Flash resetting the floor, Anthropic and OpenAI both dropping per-token rates repeatedly. Ultrafast is the first serious move to compete on latency instead: OpenAI’s own comparisons put it well ahead of Claude’s fast-mode tiers on response time for equivalent tasks. For anything built around a live, conversational AI feature — not a batch job running overnight, but something a user is sitting in front of waiting on — the model that answers first starts to matter as much as the model that answers best.
So what
If a product you’re planning leans on an AI feature that needs to feel instant — live coding assistance, a real-time support agent, anything where a two-second pause reads as broken — model selection now has a genuine speed dimension worth weighing alongside cost and quality, not an afterthought. Our AI-assisted development work includes picking the right model and hosting setup for how a feature actually needs to feel in use. Get in touch if you want help thinking through that trade-off for your build.