SGLang released day-0 production serving for open-weights GLM-5.3, demonstrating 537.6 tok/s/user under NVFP4 on 8x B300 and multi-platform optimization for NVIDIA Blackwell and AMD MI355X.

Key Takeaways

  • Day-0 production deployment support for 320B-A18B hybrid-attention architectures immediately upon weights release;
  • Optimized NVFP4/FP8 kernels sustain over 500 tokens/second per user under authentic agentic multi-turn traces;
  • Unified serving stack across NVIDIA Blackwell/Hopper and AMD Instinct accelerators eliminates vendor lock-in.