vLLM 0.28.0 landed with 584 commits, delivering stack-wide Kimi-K3 optimizations, DeepSeek-V4 sparse MLA support, adaptive speculative decoding with 65% TTFT gains, and Model Runner V2 disaggregation.
Key Takeaways
- βAdaptive speculative decoding with DFlash2/DSpark cuts time-to-first-token by 55-65% across deployments;
- βNative full-stack acceleration for cutting-edge architectures including Kimi-K3 and DeepSeek-V4 sparse MLA;
- βBroad hardware enablement across NVIDIA SM12x, AMD ROCm fused kernels, and Intel XPU backends.