OVERWRLD
News

Stranded watts are idle GPUs

A joint NVIDIA–Nscale evaluation (27 Sep 2026) ran Kimi K2.5 on GB300 NVL72 under a fixed 264.4 kW budget: DSX MaxLPS lifted managed GPUs from 140 to 192 and aggregate throughput 49.2%, with a 17% P99 TTFT trade-off.

OVERWRLD
  • AI infrastructure
  • energy
  • GPU clusters

Every unused watt inside an approved power budget is a GPU that never came online.

On 27 September 2026, Sarah McKenney and Harry Petty published How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency, a measured walkthrough of policy-governed power sharing on real inference workloads. The joint evaluation with Nscale ran Kimi K2.5 (FP4, NVIDIA Dynamo and TensorRT-LLM) on NVIDIA GB300 NVL72 systems at Nscale’s Verne campus in Keflavík, Iceland. Both configurations stayed under the same 264.4 kW provisioned power budget.

What static provisioning leaves on the table

AI factories are planned for the rare moment when every GPU hits peak draw at once. Training and inference rarely sit there. Prefill, decode, collectives, checkpoints, and idle gaps leave headroom inside each node’s reservation that a neighbor cannot touch. The facility stays under its limit while additional GPUs stay offline.

NVIDIA DSX MaxLPS monitors GPU, node, rack, and group power and reallocates within operator-defined policies. It does not raise the site’s supply. It turns variability into managed capacity under the same approved ceiling. NVIDIA also projects up to 40% more GPUs within a fixed budget for future Vera Rubin NVL72 factories; that projection is separate from the measured GB300 result below.

What the Nscale numbers show

MetricStatic baselineDSX MaxLPSChange
Managed GPUs140192+37.1%
Aggregate throughput1,084,503 tokens/s1,618,443 tokens/s+49.2%
High-throughput output per instance59,153 tokens/s59,220 tokens/s+0.1%
Low-latency output per instance2,265 tokens/s2,265 tokens/s0%
Power-budget utilization62.9%75.2%+12.3 pp
Throughput per provisioned watt4.10 tokens/s/W6.12 tokens/s/W+49.2%

The baseline used 35 four-GPU nodes (140 GPUs): two 52-GPU high-throughput instances and one 36-GPU low-latency instance. MaxLPS used 48 nodes (192 GPUs) and added a third high-throughput instance. Per-instance throughput stayed effectively flat. The gain was more concurrent work under the same reservation, not the same jobs running much faster.

Latency was not free. Median and P75 stayed within 5% of baseline. P99 time to first token rose 17% from a 15.7-second baseline. That is the judgment call for operators with strict tail SLOs: capacity and tail latency have to be accepted together.

Why this is a capacity story

This is the power half of the same argument as Schedulers need a map of the fabric and AI infrastructure still pays for the byte. Topology-aware placement keeps GPUs off congested links. Lossless compression shrinks the bytes that still have to move. Dynamic power allocation packs more useful compute under the megawatts you already purchased. None of those levers retires a substation. Each changes how much work that substation can carry.

The post’s own caveats matter. Headroom depends on complementary workload power profiles; jobs that peak together leave less to share. Control quality depends on reliable telemetry. Site electrical, cooling, and floor space still have to host the denser fleet MaxLPS unlocks. NVIDIA describes a five-stage validation path — define the managed boundary, baseline under static provisioning, introduce policy conservatively, add capacity in stages, then set production limits — before treating the configuration as production-ready.

The next halls will still buy megawatts. The open question is how many of those watts stay stranded as protective buffers instead of tokens.

Sources

  1. How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency — NVIDIA Developer Blog (27 September 2026)
  2. From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production — NVIDIA Blog