OpenAI announced its first custom inference chip Jalapeño, designed to deliver higher throughput and responsive Codex coding agent sessions.

Key Takeaways

  • Tailored for frontier inference workloads, cutting TTFT and boosting concurrent long-context sessions
  • Directly powers responsive Codex coding sessions and autonomous agent refactoring loops
  • Deployment across OpenAI enterprise compute infrastructure scheduled by year-end