vLLM 0.28.0 landed with 584 commits, delivering stack-wide Kimi-K3 optimizations, DeepSeek-V4 sparse MLA support, adaptive speculative decoding with 65% TTFT gains, and Model Runner V2 disaggregation.

Key Takeaways

  • βœ“Adaptive speculative decoding with DFlash2/DSpark cuts time-to-first-token by 55-65% across deployments;
  • βœ“Native full-stack acceleration for cutting-edge architectures including Kimi-K3 and DeepSeek-V4 sparse MLA;
  • βœ“Broad hardware enablement across NVIDIA SM12x, AMD ROCm fused kernels, and Intel XPU backends.