Close Menu
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    TechtroduceTechtroduce
    Subscribe
    • NEWS
    • GUIDES
    • COMPARISONS
    • REVIEWS
    TechtroduceTechtroduce
    Home » NVIDIA Vera Rubin NVL72 Debuts in MLPerf With 3.7x Gain
    NEWS Updated:September 17, 2026

    NVIDIA Vera Rubin NVL72 Debuts in MLPerf With 3.7x Gain

    Abyan KhanBy Abyan KhanSeptember 17, 2026Updated:September 17, 2026No Comments4 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Vera Rubin NVL72 rack photographed in a data-center environment with blue lighting
    Share
    Facebook Twitter LinkedIn Pinterest Email

    NVIDIA has published the first MLPerf Inference results for its Vera Rubin NVL72 platform, claiming up to 3.7 times the throughput of its current GB300 NVL72 system. The results were released as part of MLPerf Inference v6.1 on September 16 and cover demanding AI workloads including Qwen3-VL and DeepSeek-R1. NVIDIA submitted the Vera Rubin figures as preview results, providing an early standardized look at how its next-generation rack-scale platform compares with Blackwell Ultra.

    The largest gain came in Qwen3-VL, where NVIDIA says a Vera Rubin NVL72 rack delivered up to 3.7x higher throughput than GB300 NVL72 across MLPerf’s offline, server and interactive scenarios. The submission used vLLM alongside NVIDIA’s open-source Dynamo inference framework. On DeepSeek-R1, Vera Rubin NVL72 reached up to 2.5x the throughput of GB300 NVL72 using TensorRT-LLM.

    NVIDIA attributes the gains to changes across both hardware and software rather than the Rubin GPUs alone. The platform uses enhanced Tensor Cores and Transformer Engine capabilities to accelerate inference’s prefill and decode stages, while NVFP4 precision is used to reduce the memory footprint of model weights, attention operations and KV cache. NVIDIA also relied heavily on disaggregated serving, separating prefill and decoding workloads while using large-scale expert parallelism for mixture-of-experts models.

    Official NVIDIA Vera Rubin NVL72 rack render shown front-on against a dark background

    The rack’s communication architecture is another major part of the design. Vera Rubin NVL72 uses sixth-generation NVLink and NVLink Switch technology, which NVIDIA says delivers 10x higher packet rates and three times lower latency than off-the-shelf Ethernet for the scale-up domain. Those interconnect improvements are intended to let 72 GPUs behave more efficiently as one tightly connected system when inference workloads are distributed across the rack.

    The MLPerf announcement follows other recent performance claims for Rubin, including NVIDIA’s Vera Rubin and DSX efficiency testing, where the company has increasingly focused on metrics such as tokens per megawatt and cost per token rather than raw accelerator performance alone. NVIDIA argues that inference economics are determined by how many useful tokens a system can produce within power, networking and infrastructure constraints. That emphasis becomes increasingly important as AI companies deploy reasoning and agentic models that can generate substantially more tokens per task than conventional chat workloads.

    MLPerf v6.1 also provided new scaling data for the existing GB300 NVL72 platform. NVIDIA submitted a DeepSeek-R1 configuration using four GB300 NVL72 racks, totaling 288 GPUs, and reported 99% scaling efficiency in the offline scenario. In practical terms, throughput increased almost proportionally as the deployment expanded from one rack to four, indicating that networking and workload orchestration avoided most of the scaling losses that can appear in larger multi-rack systems.

    GB300 NVL72 also posted results in MLPerf’s WAN 2.2 text-to-video workload. NVIDIA says the rack-scale configuration produced 0.65 720p videos per second at 5.7 seconds per video, representing nine times the throughput and 7.5 times lower latency than a single-node configuration. The company is positioning that ability to scale across racks as another important part of AI-factory economics, alongside the hardware improvements arriving with Rubin.

    Software optimization contributed additional gains even without a new GPU generation. NVIDIA says GB300 NVL72 performance on Qwen3-VL improved by as much as 1.6x compared with its MLPerf Inference v6.0 submission, helped by lower-precision KV cache, kernel fusion, improved kernels and disaggregated serving. NVIDIA also says post-submission software work has produced further performance improvements on GPT-OSS-120B and DLRMv3, although those later figures have not yet been verified by MLCommons.

    The results come as NVIDIA broadens its focus from individual accelerator benchmarks toward whole-data-center efficiency. Projects such as the AI energy management alliance involving NVIDIA and Google reflect the growing importance of power availability as AI infrastructure scales. For operators deploying hundreds or thousands of accelerators, improvements in rack efficiency and near-linear scaling can matter as much as the performance of an individual GPU.

    Vera Rubin NVL72’s first MLPerf appearance therefore provides an early standardized benchmark for NVIDIA’s next platform, but the figures should still be viewed as preview results rather than the final performance ceiling. NVIDIA says continuing software optimization is expected to raise Rubin performance further as the platform matures. For now, the headline result is a claimed 3.7x throughput improvement over GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, alongside 99% multi-rack scaling efficiency demonstrated by Blackwell Ultra.

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Abyan Khan
    • Instagram
    • LinkedIn

    Abyan Khan is a dedicated writer and tech enthusiast currently pursuing a Bachelor’s degree in Information Technology. With over 3 years of professional writing experience, he specializes in crafting clear, engaging, and informative content across a range of topics, particularly in the tech and gaming industries. Abyan combines his academic knowledge with real-world insights to deliver articles that are both well-researched and reader-friendly.

    Related Posts

    Fortnite Leak Reveals Persona 5 and Crash Bandicoot Sprites

    September 17, 2026

    Monster Hunter Wilds: Ascendance Reveals Teostra and Gameplay Changes

    September 17, 2026

    PRAGMATA Gets Free Mega Man Pack DLC on September 17

    September 17, 2026
    Leave A Reply Cancel Reply

    Google Techtroduce

    See more Techtroduce stories on Google.

    Add us on Google

    • Fortnite Leak Reveals Persona 5 and Crash Bandicoot SpritesSeptember 17, 2026
      Fortnite dataminers have uncovered Persona 5 and Crash Bandicoot Sprites in the latest update, expanding Chapter 7 Season 4's crossover roster.
    • NVIDIA Vera Rubin NVL72 Debuts in MLPerf With 3.7x GainSeptember 17, 2026
      NVIDIA says Vera Rubin NVL72 delivered up to 3.7x the inference throughput of GB300 NVL72 in its first MLPerf Inference v6.1 preview results.
    • Monster Hunter Wilds: Ascendance Reveals Teostra and Gameplay ChangesSeptember 17, 2026
      Capcom reveals Teostra for Monster Hunter Wilds: Ascendance alongside faster weapon swaps, equipment changes outside tents and other improvements.
    TRENDING NOW

    Dragon Quest VII: Reimagined Demo Shows Big Switch 2 Performance Boost

    Delta Force: Black Hawk Co-Op Free Campaign

    GTA 4 Set To Respawn On PS5 And Xbox Series X|S This Year

    Fortnite v42.10 Brings Back Overwatch and Mortal Kombat Weapons

    Facebook Instagram YouTube
    © 2026 Techtroduce. All Rights Reserved | Cookies Policy | Privacy Policy | Contact Us | About Us | Corrections Policy

    Type above and press Enter to search. Press Esc to cancel.

    Manage Consent
    To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
    Functional Always active
    The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
    Preferences
    The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
    Statistics
    The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
    Marketing
    The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
    • Manage options
    • Manage services
    • Manage {vendor_count} vendors
    • Read more about these purposes
    View preferences
    • {title}
    • {title}
    • {title}