OpenAI has previewed Ultrafast, a new API service tier built to make GPT-5.6 Sol significantly faster for developers and businesses. The company says the tier can run the model at up to 14× the speed, marking a major step toward more responsive AI experiences.
Powered by Cerebras, Ultrafast can deliver up to 750 output tokens per second. That kind of throughput could make AI systems feel far more immediate in practical use cases, from live coding copilots to real-time customer support and interactive creative tools.
Why it matters
Speed is one of the biggest factors shaping how useful AI feels in everyday products. By reducing wait times, faster inference can help developers build AI applications that are smoother, more conversational, and better suited to time-sensitive workflows.
- For developers: faster responses can improve product design and user engagement.
- For businesses: lower latency can support higher-volume, real-time AI services.
- For users: AI tools can feel more natural, responsive, and productive.