vLLM v0.31.0 (717 commits, 307 contributors, 96 new) adds the vllm preload weight-cache daemon for fast engine restarts and experimental CRIU engine snapshots, makes FlashMLA mega attention with NVFP4 compressed KV the SM100 default for DeepSeek-V4.1-Flash, brings draft-model speculative decoding to Model Runner V2, gates per-request multimodal kwargs behind a flag, and fixes prefix-cache key collisions between LoRA names and cache_salt. Several breaking changes ship too.

Key Takeaways

  • ✓Scale: 717 commits from 307 contributors (96 new).
  • ✓Fast restart: vllm preload keeps post-quantized weights resident in GPU memory across restarts (DP, MTP drafts, /health); experimental CRIU snapshots restore an initialized TP1 engine.
  • ✓Security: per-request multimodal kwargs rejected unless --trust-request-mm-kwargs; LoRA path now part of the block hash to stop prefix-cache collisions.
  • ✓Breaking: tokenizer_mode="slow" removed; online quantization="fp8" replaced by fp8_per_tensor; --enforce-eager also disables JIT warmup.
  • ✓Model paths: GLM-5.3-Flash metadata ops 1.6–4.8x faster and 3 GiB indexer workspace saved; fused KimiViT QK RoPE up to 29x; ~25 s faster Qwen3.8-Flash-Next weight loading on DGX Spark.
🧭

Not enough VRAM? Compare cloud API and self-hosting costs

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

vLLM v0.31.0 (717 commits, 307 contributors) targets restart cost and multi-tenant safety. The new vllm preload daemon keeps post-quantized weights resident in GPU memory across engine restarts (with DP, MTP drafts, a /health endpoint and readiness wait), and experimental vllm snapshot create/restore uses CRIU to restore an initialized TP1 engine. DeepSeek-V4.1-Flash now defaults to FlashMLA mega attention with NVFP4 compressed KV on SM100; Model Runner V2 gains draft-model speculative decoding and custom logits processors; --max-num-active-seqs caps running admission. Security: per-request multimodal kwargs are rejected unless --trust-request-mm-kwargs is set, and LoRA names/paths can no longer collide with cache_salt in prefix-cache keys. No end-to-end throughput headline was published; component numbers include 1.6–4.8x faster GLM-5.3-Flash metadata ops, up to 29x fused KimiViT RoPE, and ~25 s faster Qwen3.8-Flash-Next loading on DGX Spark. Breaking changes: tokenizer_mode="slow" removed, online FP8 via fp8_per_tensor, renamed Mamba prefix-cache flag, --enforce-eager also disables JIT warmup. Install with pip install vllm or vllm/vllm-openai:v0.31.0.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.