Back to news
launchOpenAICerebras2026-08-13

OpenAI previews GPT-5.6 Sol Ultrafast tier powered by Cerebras at up to 750 tok/s and 14x faster

OpenAI began previewing the Ultrafast service tier for GPT-5.6 Sol on August 13. Powered by Cerebras WSE chips, it generates up to 750 output tokens per second, up to 14x faster than standard.

On August 13, OpenAI began previewing a new service tier called Ultrafast for GPT-5.6 Sol. Powered by Cerebras wafer-scale chips, it generates up to 750 output tokens per second, about 14 times faster than standard processing. Ultrafast is distinct from the existing Fast tier, which only delivers 2.5x speed at twice the price. Pricing for Ultrafast has not been published, but OpenAI indicated the speed gain will not come with a price premium.

OpenAI said Ultrafast will roll out first in the API to a select group of customers, with access expanding as capacity grows. Target scenarios include incident response (production outage diagnostics), financial research and fraud detection, conversational customer support, real-time commerce Q&A, and interactive research workflows where response time matters.

On Cerebras' Humanity Last Exam benchmark, GPT-5.6 Sol Ultrafast completed all 2,500 questions in 11 hours 11 minutes, compared to 78 hours 27 minutes for Claude Fable 5 at comparable accuracy, a roughly 7x speedup. On the GDP-Val economically valuable knowledge work benchmark, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation. Cerebras put the model at 11x faster than Claude Fable 5 and 5x faster than Opus 4.8 Fast mode.

GPT-5.6SolUltrafastCerebrasOpenAI