On Oct 9 the vLLM team, with Inferact, Red Hat and NVIDIA, shipped early Vera Rubin NVL72 support: CUDA 13.4 nightly images already serve DeepSeek, Kimi, GLM and MiniMax. Vendor-reported results: up to 7.84x per-GPU throughput vs GB200 on SemiAnalysis AgentX (MiniMax M3, matched interactivity) and up to 3.7x vs GB300 NVL72 on the MLPerf Inference v6.1 VLM benchmark (Qwen3-VL-235B-A22B).
Key Takeaways
- ✓Versus GB200 NVL72: 5x NVFP4 FLOPS, ~2.4x HBM4 bandwidth, 1.7x bidirectional NVLink, 2–4x faster softmax exponentials (vLLM blog)
- ✓AgentX with MiniMax M3: up to 7.84x per-GPU throughput at matched interactivity, 5.18x under a 150 TPS cap (vendor-reported, early)
- ✓MLPerf Inference v6.1: Qwen3-VL-235B-A22B on vLLM + Dynamo, up to 3.7x vs GB300 NVL72
- ✓CUDA 13.4 locality domains split MoE weights for ~1.2x average MoE-layer speedup in small-token decode
- ✓Image:
vllm/vllm-openai:cu134-nightly(CUDA 13.4 + PyTorch 2.15)

Key Decision Metrics at a Glance
Not enough VRAM? Compare cloud API and self-hosting costs
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Background
Agentic serving is bound by both throughput and interactivity; on GB200, MoE decode is limited by HBM weight reads and expert-parallel all-to-all traffic. Whether open-source engines run well on day one of NVIDIA's Vera Rubin matters to every deployer.
How it works
Per the vLLM blog, Rubin (sm107) is in the Blackwell family, so vLLM's sm100f kernels run unmodified. vLLM adds CUDA 13.4 locality domains that split MoE weights column-wise so each domain's SMs read only local HBM (~1.2x average MoE-layer speedup in small-token decode), Rubin-tuned FlashInfer 0.7.0 attention/GEMM/MoE kernels, and an FP8 MiniMax Sparse Attention prefill kernel.
Benchmarks (vendor-reported, early)
On SemiAnalysis AgentX with MiniMax M3: up to 7.84x per-GPU throughput vs GB200 at matched interactivity and 5.18x under a 150 TPS constraint. On MLPerf Inference v6.1 VLM (Qwen3-VL-235B-A22B, vLLM + Dynamo): up to 3.7x vs GB300 NVL72. Hardware: 5x NVFP4 FLOPS and ~2.4x HBM bandwidth vs GB200 NVL72.
Getting started
Pull vllm/vllm-openai:cu134-nightly (CUDA 13.4, PyTorch 2.15). Use --moe-backend flashinfer_cutedsl for NVFP4 MoE and --kv-cache-dtype fp8 with --attention-backend FLASHINFER for FP8 attention. Roadmap: MegaMoE, Kimi K3 MLA/KDA kernels, DeepSeek-V4.1-Flash Rubin kernels.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.