OpenAI previewed Ultrafast mode for GPT-5.6 Sol, powered by Cerebras, generating up to 750 tokens per second—as much as 14x faster. It launches first in the OpenAI API for a select group of customers, targeting latency-sensitive workflows such as real-time voice, commerce, coding, financial research, and security response.

Key Takeaways

  • GPT-5.6 Sol Ultrafast hits up to 750 tokens/sec, about 14x faster
  • Inference is powered by Cerebras for latency-critical products
  • Rolling out first via the OpenAI API for coding, voice, commerce, and security work
ADSponsored