Skip to content
AI.info

The Pulse

NVIDIA Reports First DSX Power-Management Validation

NVIDIA says Lambda’s first deployment validation of DSX MaxLPS delivered 24% more token throughput on the same power budget. The test used a 19-node HGX B200 cluster and improved performance per watt by 23%.

NVIDIA Reports First DSX Power-Management Validation

AI.info Team ·

NVIDIA says cloud provider Lambda has completed the first deployment validation of its DSX MaxLPS power-management software, reporting 24% higher token throughput without increasing the facility’s power budget. The test ran across five racks and 19 NVIDIA HGX B200 GPU servers, comparing 19 nodes operating under managed power limits with 16 nodes running at full power.

Lambda’s results were published on September 15, 2026, as NVIDIA positioned power efficiency as a defining measure for large AI facilities. The company says the cluster increased throughput from roughly 4 million tokens per second to 5 million and improved performance per watt by 23%.

“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” Dave Ward, president of cloud services at Lambda, said in NVIDIA’s account. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.”

Lambda’s 19-node test puts a number on stranded power

DSX MaxLPS monitors power use at the GPU and rack levels, then reallocates available headroom across nodes according to workload type. NVIDIA says the approach is designed for facilities running both training and inference, where demand patterns can differ and static power provisioning may leave capacity unused.

Lambda operated 19 nodes within the same power budget used by 16 nodes at full power. NVIDIA reports that the arrangement raised cluster-wide token throughput by 24%, from approximately 4 million to 5 million tokens per second, while increasing performance per watt by 23%.

The result is a deployment measurement from one five-rack cluster, not a claim that every AI facility will achieve the same gain. NVIDIA presents the test as the first validation of DSX MaxLPS on HGX B200 GPU servers and separately projects that the software could support up to 40% more GPU capacity in suitable Vera Rubin NVL72 facilities operating under the same megawatt budget.

DSX MaxLPS sits inside a wider factory design

NVIDIA describes DSX MaxLPS as one part of a broader DSX platform for designing, simulating and operating AI factories. The platform includes DSX Sim for modelling infrastructure before deployment, DSX OS for operating workloads and facility systems, DSX Exchange for coordinating IT and operational technology, and DSX Flex for responding to utility and grid signals.

That structure matters because power allocation is only one source of wasted capacity. Network delays, cooling overhead, scheduling gaps and conservative rack provisioning can all reduce the amount of useful computation delivered by a fixed electrical supply. NVIDIA’s DSX documentation describes MaxLPS as a suite of chip, thermal, system and software technologies intended to improve performance per watt inside a fixed power envelope.

NVIDIA founder and CEO Jensen Huang has framed the constraint in simple terms: “A one-gigawatt factory will never become a two-gigawatt factory.” The company’s answer is to increase the amount of useful work produced by the power already available rather than treating additional grid capacity as the only route to growth.

From rack allocation to grid response

NVIDIA’s announcement also points to a separate deployment involving Emerald AI’s Conductor platform at an AI factory connected to Silicon Valley Power’s Flexible Load Interconnect Program. When the utility sends a demand signal, Conductor reduces consumption by slowing or rescheduling lower-priority workloads while keeping higher-priority services active.

NVIDIA says the facility automatically reduced power from four megawatts to three in under a minute, without an operator intervening or interrupting critical inference jobs. Silicon Valley Power has since sent more than 200 demand signals to the site, and NVIDIA says the system responded successfully each time.

The Santa Clara installation is not a DSX Flex deployment. NVIDIA describes it as an early commercial demonstration of the operating model that DSX Flex is intended to support, with Emerald AI Conductor expected to integrate into the platform as it matures.

Manassas will be the first dedicated DSX Flex deployment

NVIDIA says the first dedicated commercial DSX Flex deployment will be at its AI Factory Research Center in Manassas, Virginia. The planned facility is described as a 96-megawatt Vera Rubin AI factory and will build on five earlier demonstrations across two continents.

DSX Flex is designed to receive signals such as load-shedding requests, demand-response events and pricing changes, then adjust workload priorities in response. The goal is to make an AI facility flexible enough to reduce demand when the grid is constrained while preserving the services operators have marked as highest priority.

That capability gives the power-management results two distinct applications. MaxLPS reallocates capacity inside the facility to produce more tokens under a fixed budget, while DSX Flex connects workload decisions to external grid conditions. Both treat the data center’s electrical limit as an operating parameter rather than a fixed obstacle.

NVIDIA is also preparing an 800-volt power path

The company says DSX reference designs are incorporating an 800-volt direct-current architecture for denser accelerated-computing racks. NVIDIA argues that the higher-voltage design can reduce conversion complexity and improve power delivery efficiency as rack demand rises.

NVIDIA projects a 3% to 5% end-to-end efficiency gain compared with 54-volt distribution and says the architecture will be available with Vera Rubin NVL72 systems in 2027. That projection is separate from Lambda’s measured DSX MaxLPS result.

Lambda’s validation therefore establishes a concrete result for one current configuration: 19 HGX B200 nodes delivered 24% more token throughput than a 16-node full-power baseline within the same facility power budget. NVIDIA’s broader DSX program extends that idea from rack-level allocation to grid response and future power-delivery designs.

Source

NVIDIA Blog

Explore

More articles