NVIDIA has published the first MLPerf Inference results for its Vera Rubin NVL72 platform, claiming up to 3.7 times the throughput of its current GB300 NVL72 system. The results were released as part of MLPerf Inference v6.1 on September 16 and cover demanding AI workloads including Qwen3-VL and DeepSeek-R1. NVIDIA submitted the Vera Rubin figures as preview results, providing an early standardized look at how its next-generation rack-scale platform compares with Blackwell Ultra.
The largest gain came in Qwen3-VL, where NVIDIA says a Vera Rubin NVL72 rack delivered up to 3.7x higher throughput than GB300 NVL72 across MLPerf’s offline, server and interactive scenarios. The submission used vLLM alongside NVIDIA’s open-source Dynamo inference framework. On DeepSeek-R1, Vera Rubin NVL72 reached up to 2.5x the throughput of GB300 NVL72 using TensorRT-LLM.
NVIDIA attributes the gains to changes across both hardware and software rather than the Rubin GPUs alone. The platform uses enhanced Tensor Cores and Transformer Engine capabilities to accelerate inference’s prefill and decode stages, while NVFP4 precision is used to reduce the memory footprint of model weights, attention operations and KV cache. NVIDIA also relied heavily on disaggregated serving, separating prefill and decoding workloads while using large-scale expert parallelism for mixture-of-experts models.

The rack’s communication architecture is another major part of the design. Vera Rubin NVL72 uses sixth-generation NVLink and NVLink Switch technology, which NVIDIA says delivers 10x higher packet rates and three times lower latency than off-the-shelf Ethernet for the scale-up domain. Those interconnect improvements are intended to let 72 GPUs behave more efficiently as one tightly connected system when inference workloads are distributed across the rack.
The MLPerf announcement follows other recent performance claims for Rubin, including NVIDIA’s Vera Rubin and DSX efficiency testing, where the company has increasingly focused on metrics such as tokens per megawatt and cost per token rather than raw accelerator performance alone. NVIDIA argues that inference economics are determined by how many useful tokens a system can produce within power, networking and infrastructure constraints. That emphasis becomes increasingly important as AI companies deploy reasoning and agentic models that can generate substantially more tokens per task than conventional chat workloads.
MLPerf v6.1 also provided new scaling data for the existing GB300 NVL72 platform. NVIDIA submitted a DeepSeek-R1 configuration using four GB300 NVL72 racks, totaling 288 GPUs, and reported 99% scaling efficiency in the offline scenario. In practical terms, throughput increased almost proportionally as the deployment expanded from one rack to four, indicating that networking and workload orchestration avoided most of the scaling losses that can appear in larger multi-rack systems.
GB300 NVL72 also posted results in MLPerf’s WAN 2.2 text-to-video workload. NVIDIA says the rack-scale configuration produced 0.65 720p videos per second at 5.7 seconds per video, representing nine times the throughput and 7.5 times lower latency than a single-node configuration. The company is positioning that ability to scale across racks as another important part of AI-factory economics, alongside the hardware improvements arriving with Rubin.
Software optimization contributed additional gains even without a new GPU generation. NVIDIA says GB300 NVL72 performance on Qwen3-VL improved by as much as 1.6x compared with its MLPerf Inference v6.0 submission, helped by lower-precision KV cache, kernel fusion, improved kernels and disaggregated serving. NVIDIA also says post-submission software work has produced further performance improvements on GPT-OSS-120B and DLRMv3, although those later figures have not yet been verified by MLCommons.
The results come as NVIDIA broadens its focus from individual accelerator benchmarks toward whole-data-center efficiency. Projects such as the AI energy management alliance involving NVIDIA and Google reflect the growing importance of power availability as AI infrastructure scales. For operators deploying hundreds or thousands of accelerators, improvements in rack efficiency and near-linear scaling can matter as much as the performance of an individual GPU.
Vera Rubin NVL72’s first MLPerf appearance therefore provides an early standardized benchmark for NVIDIA’s next platform, but the figures should still be viewed as preview results rather than the final performance ceiling. NVIDIA says continuing software optimization is expected to raise Rubin performance further as the platform matures. For now, the headline result is a claimed 3.7x throughput improvement over GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, alongside 99% multi-rack scaling efficiency demonstrated by Blackwell Ultra.

