OpenAI previewed Ultrafast mode in the API for select developer teams. Powered by Cerebras hardware, it delivers up to ~750 tokens/sec (up to 14x faster generation). Target use cases include realtime voice, high-throughput coding and design, and automated security incident response.

Key Takeaways

  • Powered by Cerebras inference hardware, output throughput peaks at ~750 tokens/s.
  • Brings frontier-class intelligence into latency-critical real-time coding workflows.
  • Coding, design generation, and security response are primary target workloads.
ADSponsored