OpenAI previewed Ultrafast mode in the API for select developer teams. Powered by Cerebras hardware, it delivers up to ~750 tokens/sec (up to 14x faster generation). Target use cases include realtime voice, high-throughput coding and design, and automated security incident response.
Key Takeaways
- ✓Powered by Cerebras inference hardware, output throughput peaks at ~750 tokens/s.
- ✓Brings frontier-class intelligence into latency-critical real-time coding workflows.
- ✓Coding, design generation, and security response are primary target workloads.