OpenAI released InferenceX benchmarks for Jalapeño, its first custom inference silicon, showing 1.5–1.9x higher work per watt and 1.7–3.6x lower end-to-end latency across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
Key Takeaways
- ✓OpenAI officially expands into first-party silicon, publishing public model-agnostic benchmarks across open weights;
- ✓Designed around agentic serving, unified architecture handles both prefill and decode stages simultaneously;
- ✓AI-assisted chip development achieved tapeout in 9 months and enabled rapid Codex-driven model porting in 2 months.