DeepSeek previewed next-gen V4 architecture featuring sparse MoE, dropping long-context coding token cost to $0.07/1M.
Key Takeaways
- ✓Evolved Multi-head Latent Attention reduces KV cache memory footprint by 60%
- ✓Logical error rates in long-range refactoring reduced by 42% over V3
- ✓Full open-weights release guaranteed with permissive commercial license