DeepSeek previewed next-gen V4 architecture featuring sparse MoE, dropping long-context coding token cost to $0.07/1M.

Key Takeaways

  • Evolved Multi-head Latent Attention reduces KV cache memory footprint by 60%
  • Logical error rates in long-range refactoring reduced by 42% over V3
  • Full open-weights release guaranteed with permissive commercial license