- Start
- Sep 16, 202695% CONFIDENCEfrom mlcommons.org
NVIDIA Vera Rubin NVL72 Debuts in MLPerf Inference v6.1
- MLCommons published the MLPerf Inference v6.1 results on 16 September 2026, the first peer-reviewed numbers for NVIDIA's Vera Rubin NVL72, which was entered in the preview category[3][4]
- On Qwen3-VL, Vera Rubin NVL72 delivered up to 3.7x the throughput of GB300 NVL72 across offline, server and interactive scenarios using vLLM with NVIDIA Dynamo, and up to 2.5x on DeepSeek-R1 using TensorRT-LLM[1]
- Nebius, the only other Vera Rubin submitter, ran DeepSeek R1 on 36 GPUs across nine nodes and posted 16,427 tokens/s per GPU in the server scenario[6]
- A 288-GPU GB300 NVL72 submission across four racks reached 99% scaling efficiency on DeepSeek-R1, and software work lifted GB300 NVL72 on Qwen3-VL by up to 1.6x over v6.0[1]
- The round drew a record 30 submitting organizations and added End-to-End RAG and Edge Agentic Inference tests; the best DeepSeek R1 per-accelerator server result was 5.7x a year earlier[4]
- NVIDIA's claim of 30x over GB300 NVL72 on SemiAnalysis AgentX comes from preview testing outside MLPerf and is not verified by MLCommons[1][3]
- Results were published at 11:00 a.m. EDT through MLCommons' GlobeNewswire release[5]
Notable features
- One rack holds 72 Rubin GPUs and 36 Vera CPUs, with ConnectX-9 SuperNICs and BlueField-4 DPUs[2]
- 20.7 TB of HBM4 at 1,400 TB/s, 3,168 custom Olympus CPU cores and up to 54 TB of LPDDR5X per rack[2]
- Sixth-generation NVLink and NVLink Switch, which NVIDIA says gives 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet[1]
- NVFP4 precision and enhanced Tensor Cores and Transformer Engine speed both prefill and decode[1]
- NVIDIA rates it at one-tenth the cost per million tokens of GB200 NVL72[2]
References 686% CONFIDENCE
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inferenceblogs.nvidia.com· Posted Sep 23, 2026· Starts Sep 16, 2026· 17% of score
MLCommons[4] published the MLPerf Inference v6.1 results on 16 September 2026 - its GlobeNewswire release is datelined 'SAN FRANCISCO, Sept. 16, 2026' and stamped '11:00 AM EDT' - and NVIDIA's[2] post of the same day says the results were 'released today'.
- [2]75%NVIDIA Vera Rubin NVL72nvidia.com· Added Sep 23, 2026· 17% of score
NVIDIA's[1] product page for the benchmarked system: 72 Rubin GPUs and 36 Vera CPUs per rack, 20.7 TB of HBM4 at 1,400 TB/s, 3,168 Olympus CPU cores, sixth-generation NVLink, and one-tenth the cost per million tokens of GB200 NVL72.
- [3]80%MLPerf v6.1 Puts First Peer-Reviewed Numbers on Vera Rubin: Software Gains Beat New Hardwaretechtimes.com· Published Sep 17, 2026· 17% of score
Independent analysis of the round 'published September 16, 2026': Vera Rubin NVL72 was entered in the Preview category reserved for platforms due to ship by the next round, 30 organizations produced 486 results, and NVIDIA's[1][2] 30x AgentX figure is separate from, and not verified by, MLCommons[4].
- [4]95%MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Resultsmlcommons.org· Published Sep 16, 2026· Starts Sep 16, 2026 ✓· 17% of score
MLCommons' own announcement of the v6.1 round, dated September 16, 2026: a record 30 submitting organizations, two new tests (End-to-End RAG and Edge Agentic Inference), up to 5.7x on DeepSeek R1 versus a year earlier, and 'NVIDIA[1][2] Rubin and NVIDIA Vera Rubin NVL72 is in preview' among five new processors.
- [5]90%MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results (GlobeNewswire)financialcontent.com· Published Sep 16, 2026· 17% of score
The GlobeNewswire copy of MLCommons[4]' release, datelined 'SAN FRANCISCO, Sept. 16, 2026 (GLOBE NEWSWIRE)' and stamped 11:00 AM EDT, which gives the publication time of the results.
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most.