From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
NVIDIA's Vera Rubin NVL72 system made its MLPerf Inference v6.1 debut, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x higher throughput on DeepSeek-R1.
The GB300 NVL72 achieved 99% scaling efficiency across four racks with 288 GPUs.
Software optimizations in v6.1 delivered up to 1.6x higher performance over v6.0.
From the source
Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.
blogs.nvidia.com