vLLM 0.28.0 landed with 584 commits, bringing stack-wide Kimi-K3 support, end-to-end DeepSeek-V4 sparse MLA with DFlash2/DSpark speculative decoding, Model Runner V2 disaggregation, and tiered disk KV cache.
Key Takeaways
- ✓Adaptive speculative token budgets slash end-to-end TTFT by 55–65% with 1.5–3x faster sequence-parallel kernels;
- ✓End-to-end DeepSeek-V4 sparse MLA and Kimi-K3 shared-expert sharding save ~17 GiB VRAM per GPU;
- ✓Tiered KV cache introduces a disk tier, doubles max batched tokens to 16,384, and supports diverse hardware backends.