OpenAI previewed Ultrafast mode for GPT-5.6 Sol powered by Cerebras, achieving up to 750 tokens per second (14x standard speed) targeted at real-time voice, low-latency agents, and high-throughput coding workflows.

Key Takeaways

  • Powered by Cerebras wafer-scale chips, GPT-5.6 Sol reaches up to 750 tokens/s (~14x speedup);
  • Rolling out to select API customers targeting high-throughput coding, instant code refactoring, and voice agents;
  • End-to-end full stack optimization designed specifically for latency-critical developer workflows.