NVIDIA RTX Spark announced local-agent inference optimizations: up to 1.9× higher llama.cpp throughput on GeForce RTX 5090, 1.2× vLLM on RTX PRO 6000 Blackwell, and up to 1.4× on a two-system DGX Spark cluster—another step for faster on-device and local agent serving.

Key Takeaways

  • llama.cpp: up to ~1.9× throughput on GeForce RTX 5090 for local agent inference.
  • vLLM: ~1.2× on RTX PRO 6000 Blackwell; up to ~1.4× on a two-system DGX Spark cluster.
  • Focus: faster local/edge agent serving stacks, not only cloud inference.
ADSponsored