SGLang released day-0 production serving for open-weights GLM-5.3, demonstrating 537.6 tok/s/user under NVFP4 on 8x B300 and multi-platform optimization for NVIDIA Blackwell and AMD MI355X.
Key Takeaways
- ✓Day-0 production deployment support for 320B-A18B hybrid-attention architectures immediately upon weights release;
- ✓Optimized NVFP4/FP8 kernels sustain over 500 tokens/second per user under authentic agentic multi-turn traces;
- ✓Unified serving stack across NVIDIA Blackwell/Hopper and AMD Instinct accelerators eliminates vendor lock-in.