OpenAI shared an early look at Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second. The tier is powered by Cerebras hardware and is in limited preview with a selected group of customers in coding, commerce, financial research, and support. OpenAI frames the point as removing the usual trade-off in which real-time latency forced users to drop to a smaller model, citing internal use in incident response — reading logs, analyzing traces, and preparing fixes while an outage is still unfolding — and in research, where overnight batches compress into same-day loops.





