# NVIDIA — NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

- Company: NVIDIA (nvidia.com)
- Announced: 2026-09-16T15:00:48+00:00
- Category: infrastructure-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
- Record: https://forck.live/items/11417-nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1
- Subject: AI platform
- Models affected: DeepSeek-R1, Qwen3-VL, GPT-OSS-120B, DLRMv3, Qwen3.6-27B

NVIDIA's Vera Rubin NVL72 system made its MLPerf Inference v6.1 debut, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x higher throughput on DeepSeek-R1. The GB300 NVL72 achieved 99% scaling efficiency across four racks with 288 GPUs. Software optimizations in v6.1 delivered up to 1.6x higher performance over v6.0.

## Evidence

Verbatim from https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/:

> Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.

---

Record: https://forck.live/items/11417-nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
