OpenAI previewed Ultrafast mode for GPT-5.6 Sol, powered by Cerebras, generating up to 750 tokens per second—as much as 14x faster. It launches first in the OpenAI API for a select group of customers, targeting latency-sensitive workflows such as real-time voice, commerce, coding, financial research, and security response.
Key Takeaways
- ✓GPT-5.6 Sol Ultrafast hits up to 750 tokens/sec, about 14x faster
- ✓Inference is powered by Cerebras for latency-critical products
- ✓Rolling out first via the OpenAI API for coding, voice, commerce, and security work