SGLang, a leading open-source high-throughput inference engine for large models, officially releases milestone version v0.5.21, merging 779 PRs across 227 contributors. The release introduces dynamic on-the-fly role switching between Prefill and Decode (PD) instances without restarting, an upgraded high-performance Rust core for the default prefix cache, and 22% faster first-token generation for DeepSeek-V4.1 on long prompts. It also debuts the /v1/decisions and /v1/score APIs, alongside complete support for next-gen models including MiMo-V2.6, DiffusionGemma, and GLM-5.3-Flash on AMD MI355X.

Key Takeaways

  • ✓Pioneers zero-restart dynamic role switching between prefill and decode instances to eliminate cluster capacity skew
  • ✓Upgrades default prefix cache to a native Rust core, cutting DeepSeek-V4.1 TTFT on long prompts by 22%
  • ✓Improves Kimi K3 prefill throughput by 20.6% and adds native AMD MI355X support with FP8/MXFP4 MoE kernels
🧭

Not enough VRAM? Compare cloud API and self-hosting costs

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Serving modern large language models under heavy concurrency requires navigating fundamentally divergent hardware requirements: Prefill is compute-bound, demanding raw tensor FLOPS to ingest prompt contexts, while Decode is memory-bandwidth-bound, iteratively fetching KV caches. While Prefill-Decode (PD) disaggregation addresses this asymmetry, traditional deployments assign instance roles statically. Under volatile workloads, static clusters suffer from severe capacity skew and underutilized GPUs. Concurrently, prefix caching management implemented in Python frequently creates lock contention and memory fragmentation under dense multi-tenant traffic.

架构亮点与底层机制

SGLang v0.5.21 introduces transformative architectural improvements:

  1. Zero-Restart Dynamic PD Role Switching: Instances dynamically transition between Prefill and Decode roles on the fly without service interruptions, enabling schedulers to adapt to traffic patterns in milliseconds.
  2. Native Rust-Core Prefix Cache: The Radix Tree prefix cache engine is ported to a native Rust implementation and enabled by default, eliminating Python GIL bottlenecks and delivering sub-microsecond cache lookups.
  3. Dedicated Decisions and Score APIs: Introduces /v1/decisions and /v1/score endpoints, empowering developers to utilize models as ultra-low-latency discrete classifiers and multi-candidate rerankers without autoregressive overhead.
  4. Frontier Hardware & Model Expansion: Validates GLM-5.3-Flash across AMD MI355X accelerators with native FP8/MXFP4 MoE kernels and multi-token prediction (MTP) speculative decoding.

权威 Benchmark 与实测跑分对比

Benchmarked across industrial production traces:

  1. 22% Lower TTFT on DeepSeek-V4.1: Yields a 22% reduction in time-to-first-token across 32k-128k long-context scenarios under distributed serving.
  2. 20.6% Higher Prefill Throughput on Kimi K3: Demonstrates a 20.6% throughput gain under disaggregated PD serving.
  3. Layer-Boundary Native Communication: Resolves tensor parallel discrepancies under PP, DP, and context parallelism (CP), stabilizing extended rollouts.

开发者实战落地与开箱指南

SGLang v0.5.21 is available via uv pip install --prerelease=allow sglang==0.5.21 and prebuilt Docker images on Docker Hub for CUDA 13 and ROCm 10. Platform engineering teams can deploy dynamic PD switching and the Rust cache engine to achieve over 20% aggregate throughput gains out of the box.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.