
16 Sep 2026
NVIDIA Vera Rubin NVL72 debuts on MLPerf Inference — up to 3.7x GB300
NVIDIA’s company blog said Vera Rubin NVL72’s first MLPerf Inference v6.1 preview submission delivered up to 3.7x better throughput than GB300 NVL72 on Qwen3-VL, and up to 2.5x on DeepSeek-R1, with GB300 also showing 99% multi-rack scaling efficiency.
HARDWARE desk — same-day company primary that puts Vera Rubin NVL72 on the public MLPerf Inference scoreboard against GB300, giving buyers a concrete throughput and scaling comparison for AI inference racks.
Company claim: Vera Rubin NVL72 delivered up to 3.7x better throughput than GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios. Throughput here means how many tokens or queries the rack can serve. Qwen3-VL is a vision-language model in the suite — it reads text and images. Offline, server, and interactive are the three load patterns MLPerf uses: a batch of work, a live stream of requests, and a snappier interactive stream. NVIDIA says those Qwen3-VL runs used vLLM with NVIDIA Dynamo, an open inference stack. File 3.7x and those three scenarios as NVIDIA’s, via MLPerf entries 6.1-0106 and 6.1-0074. This desk did not time a rack.
Company claim: on DeepSeek-R1, Vera Rubin NVL72 delivered up to 2.5x higher throughput than GB300 NVL72. DeepSeek-R1 is a reasoning model in the same suite. NVIDIA says those runs used TensorRT-LLM, the company’s library for serving large language models. File 2.5x as NVIDIA’s, via the same named entries. This desk did not rerun DeepSeek-R1.
Company claim: a 288-GPU, four-rack GB300 NVL72 DeepSeek-R1 submission reached 99% scaling efficiency in the offline scenario versus a single 72-GPU rack. NVL72 is NVIDIA’s 72-GPU rack-scale system — seventy-two of the company’s AI chips wired as one cabinet. Scaling efficiency means how much extra throughput you get when you add more GPUs. 99% means four racks did almost four times the work of one. File 288 GPUs, four racks, and 99% as NVIDIA’s, via MLPerf entries 6.1-0073 and 6.1-0074. This desk did not rent those racks.
Company claim: software optimizations lifted GB300 NVL72 Qwen3-VL performance up to 1.6x versus the prior v6.0 results. File 1.6x as NVIDIA’s software-stack claim on the same blog. This desk did not compare the two rounds independently. Extra kernel-fusion and cache-precision talk stays in Sources.
Company notes on the same post: Nebius also submitted Vera Rubin NVL72 preview results. NVIDIA submitted Jetson AGX Thor on a new Edge-Agentic benchmark with Qwen3.6-27B — a smaller model meant for devices at the edge of the network, not a full data-center rack. The company says 19 partners participated broadly. File Nebius, Jetson AGX Thor, and the 19-partner count as NVIDIA’s. Partner names stay in Sources. This desk did not audit partner submissions.
Plain English for the rest of the card: MLPerf = industry benchmark suite run by MLCommons. NVL72 = NVIDIA’s 72-GPU rack-scale system. Vera Rubin = NVIDIA’s next-generation platform after Blackwell / GB300. GB300 = the current NVIDIA rack Vera Rubin is being scored against. throughput = how many tokens or queries the rack can serve. inference = running a trained model to answer a query, not teaching a new one. preview submission = on the scoreboard, not yet in the available-to-buy MLPerf bucket. scaling efficiency = how much extra throughput you get when you add more GPUs. Qwen3-VL = the suite’s vision-language model. DeepSeek-R1 = the suite’s reasoning model.
PRIMARY here: NVIDIA’s 16 Sep 2026 company blog — Tier A PRIMARY company source, the original record. MLCommons’ public Inference results page is the suite home NVIDIA cites (retrieved Sep 16, 2026, Closed Division); it is not a second originating NVIDIA newsroom. The first Vera Rubin NVL72 preview submission, the up-to-3.7x Qwen3-VL throughput line, the up-to-2.5x DeepSeek-R1 line, the 288-GPU / four-rack / 99% GB300 scaling line, the up-to-1.6x GB300 Qwen3-VL software lift versus v6.0, the Nebius preview note, the Jetson AGX Thor Edge-Agentic / Qwen3.6-27B note, the 19-partner count, and the named MLPerf entries 6.1-0106, 6.1-0074, and 6.1-0073 are NVIDIA-attributed. Throughput and scaling numbers stay company-attributed to those named entries — not independently re-measured here. Extra preview-testing scores and post-submission unverified numbers stay in Sources. NOT claimed: a U.S. retail ship date, a list price, an independent lab re-run, a customer win count, that Vera Rubin is generally available, that this desk reran MLPerf, a “world’s fastest” ranking beyond NVIDIA’s wording on these named entries, a stock tip, or investment advice. Distinct from the already-filed nvidia-dropless-moe-jax, crusoe-perplexity-partnership, dmatrix-nvlink-fusion-raptor, euclyd-200m-series-a, rune-40m-series-a, and openai-chatgpt-sponsored-agents.
RELATED
- NVIDIA shows ~10× faster dropless MoE training in JAX
- Crusoe and Perplexity ink a multi-year deal for train-to-serve AI cloud
- NVIDIA blog: d-Matrix to connect Raptor XPUs via NVLink Fusion
- EUCLYD raises more than €200M Series A for efficient AI inference chips
- OpenAI tests Sponsored Agents inside ChatGPT Ads
- Rune raises $40M Series A for RELIC solar-powered modular AI compute
On 16 Sep 2026 NVIDIA’s company blog published “NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut,” by Zhihan Jiang. The post said Vera Rubin NVL72 made its first MLPerf Inference preview submission in the v6.1 suite. MLPerf is an industry benchmark suite run by MLCommons — a shared test so vendors can publish comparable AI scores. A preview submission means the system is on the scoreboard but is not yet in the “available to buy or rent” bucket. That NVIDIA blog is the filing event. These are NVIDIA’s figures, tied to named MLPerf Closed Division entries. This desk did not rerun the benchmark. “Leading” is company wording — not a desk ranking.