Back to feed
Artificial intelligenceDEVELOPING

NVIDIA presents new AI infrastructure efficiency results at AI Infra Summit

NVIDIA presented new Vera Rubin and DSX efficiency results at the AI Infra Summit, including up to 3.7x higher throughput for Vera Rubin NVL72 than GB300 NVL72 in preliminary MLPerf Inference v6.1 results. Lambda reported a 23% gain in performance per watt.

Preferred on Google
Text size

Image: blogs.nvidia.com · Author: NVIDIA Writers · source articleEditorial excerpt for news reporting

NVIDIA presented new AI infrastructure results at the AI Infra Summit in Santa Clara. Ian Buck, the company’s vice president for hyperscale and high-performance computing, discussed Vera Rubin, DSX power management and new partner integrations.

In preliminary MLPerf Inference v6.1 results, the Vera Rubin NVL72 system delivered up to 3.7x higher throughput than NVIDIA GB300 NVL72. MLPerf is operated by the MLCommons consortium, which reviews results before publication.

The event also included results from Lambda, which reported a 23% improvement in performance per watt using NVIDIA DSX MaxLPS. Lambda ran 19 nodes within the power budget normally assigned to 16 full-power nodes and raised cluster throughput from about 4 million to 5 million tokens per second.

Vera Rubin and MLPerf results

A 288-GPU submission using four GB300 NVL72 racks reached 99% scaling efficiency. NVIDIA separately said software optimizations in MLPerf Inference v6.1 produced up to 1.6x higher performance than version 6.0.

NVIDIA describes its infrastructure as a full-stack platform covering Vera Rubin systems, Dynamo and NeMo software, NVLink, Spectrum-X Ethernet and ConnectX SuperNIC networking, plus BlueField technologies for context-memory storage and infrastructure security.

DSX manages AI factory power

Emerald AI and NVIDIA, working with Silicon Valley Power, demonstrated a flexible-load program for AI factories. The system responded to hundreds of grid signals while maintaining the performance of priority AI workloads.

DSX Flex can automatically pause lower-priority jobs during load-reduction requests and resume them afterward. It can respond to load-shedding, demand-response and electricity-pricing signals.

Partnerships and additional figures

Amazon’s Annapurna Labs is working with NVIDIA on custom NVHBM high-bandwidth memory technology. d-Matrix is integrating NVLink Fusion with Vera CPUs and Raptor XPU accelerators for low-latency inference, while Pinterest is using Blackwell and Dynamo for conversational visual discovery.

NVIDIA said Vera Rubin NVL72 can support up to 40% more GPUs within the same power budget in suitable deployment environments. Combined with Groq 3 LPX, the company also reported up to 35 times higher token throughput per megawatt than GB200 NVL72 for models exceeding 2 trillion parameters at long context.

On a 100,000-token-context Qwen 3.8 27B workload, Groq 3 LPX reached 2,529 output tokens per second per user. SemiAnalysis AgentX data showed Vera Rubin NVL72 delivering up to 30 times higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model.

What we know

  • Vera Rubin NVL72 delivered up to 3.7x the throughput of GB300 NVL72 in preliminary MLPerf Inference v6.1 results.
  • Lambda improved performance per watt by 23% with DSX MaxLPS.
  • Emerald AI and NVIDIA demonstrated automated AI-factory load reduction in response to hundreds of Silicon Valley Power signals.
  • Groq 3 LPX reached 2,529 output tokens per second per user on a Qwen 3.8 27B workload.

What is being verified

  • The newsroom is checking the report that vera Rubin NVL72 delivered up to 3.7x the throughput of GB300 NVL72 in preliminary MLPerf Inference v6.1 results.
  • Reporting from NVIDIA Newsroom is being compared; a second independent confirmation is not yet available.
If a new independent confirmation or correction appears, it will be added to the story timeline automatically.
View sources1

COMMUNITY

Discussion

0

No comments yet. Start the discussion.