NVIDIA published the first on-silicon benchmark metrics for its Vera Rubin architecture on long-horizon agent workloads, demonstrating up to 30x throughput per megawatt and 35x lower token cost versus GB300 NVL72.

Key Takeaways

  • Inference benchmarks rebased from short chat to multi-step agent coding workloads with deep context expansion;
  • Benchmarked against DeepSeek V4 Pro on SemiAnalysis AgentX real coding traces as the reference workload;
  • Delivers up to 30x throughput per megawatt and 35x lower token cost compared to GB300 NVL72.