Mistral AI rolls out major API runtime optimizations for Codestral 2501 and Mistral Large 3. Featuring 256K context and specialized Fill-in-the-Middle (FIM) kernels, developers get sub-40ms time-to-first-token in IDE plugins at $0.30/$0.90 per 1M tokens.

Key Takeaways

  • Codestral 2501 introduces low-latency FIM kernels with sub-40ms time-to-first-token in IDEs;
  • 256K context window enables full-repo AST ingestion and cross-file signature completion;
  • Competitive pricing of $0.30 input / $0.90 output on La Plateforme API.