OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
GPT-5.6 Sol now runs at up to 750 output tokens/sec in an "Ultrafast" mode—roughly 14x standard speed—powered by Cerebras hardware, meaning you no longer have to drop to a smaller model to hit real-time latency for your top-tier model. This unlocks latency-sensitive agentic and interactive workflows (incident response, live support, market analysis) at frontier quality, but it's preview-only to a limited customer set gated by Cerebras capacity, so don't architect production dependencies on it yet.
GPT-5.6 Sol can now process at 14x the standard speed, generating up to 750 output tokens per second, enabling real-time applications like incident response and customer service without sacrificing model capability, and is being made available in preview to a select group of customers.