IT Brief UK - Technology news for CIOs & IT decision-makers
United Kingdom
Nvidia backs AI power efficiency push with new deals

Nvidia backs AI power efficiency push with new deals

Fri, 18th Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Nvidia outlined a series of collaborations and performance results focused on AI infrastructure efficiency, particularly how AI systems use electricity as demand for agentic AI workloads grows.

Speaking at the AI Infra Summit in Santa Clara, Ian Buck, Vice President of Hyperscale and High-Performance Computing at Nvidia, said the industry is shifting from peak chip performance to output measured against power consumption.

That shift is pushing customers and partners to treat token throughput per megawatt as a key metric for AI deployments, especially as larger models and agent-based applications put more pressure on data centre power budgets.

Power focus

Among the announcements, Emerald AI worked with Nvidia and Silicon Valley Power on a commercial AI factory flexible-load programme designed to cut electricity demand when required by the grid while keeping priority AI tasks running.

The system responded to hundreds of utility demand signals using software that can throttle lower-priority AI jobs during periods of grid stress and restore normal operations later.

Nvidia said this approach could help AI facilities act as controllable loads, potentially allowing utilities to support more computing sites without immediately expanding electricity infrastructure.

Lambda, an AI cloud provider, also released test results for Nvidia's DSX MaxLPS software on Blackwell servers. According to the figures, Lambda ran 19 nodes within a power budget usually assigned to 16 full-power nodes.

That increased cluster-wide token throughput by 24%, from about 4 million to 5 million tokens per second, while improving performance per watt by 23%.

For operators trying to expand output without securing additional power, the figures highlight a central pressure in the AI infrastructure market: electricity is becoming as important a constraint as chip supply.

Platform links

Nvidia also disclosed several platform partnerships. Amazon's Annapurna Labs is working with Nvidia on NVHBM custom high-bandwidth memory technology, while d-Matrix is integrating with NVLink Fusion to combine Nvidia Vera CPUs with d-Matrix Raptor XPUs for low-latency inference.

Pinterest is using the Nvidia Blackwell platform and Nvidia Dynamo inference software for conversational AI in visual discovery.

Nvidia said its broader product lineup spans Vera Rubin systems, Dynamo software, NeMo libraries and networking products including NVLink, Spectrum-X Ethernet, ConnectX SuperNICs, BlueField-powered storage and BlueField DPUs.

Within that portfolio, Nvidia highlighted Vera Rubin NVL72 as a central system for large-scale AI deployments. It said DSX MaxLPS can provide up to 40% more GPU capacity within the same megawatt budget in suitable environments, and up to 35% higher token throughput without requiring new power lines.

Nvidia also linked Vera Rubin to Groq 3 LPX for inference workloads, saying the combined setup can deliver up to 35 times higher token throughput per megawatt than GB200 NVL72 for very large models with long context windows.

On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX reached 2,529 output tokens per second per user, according to Nvidia.

Benchmark claims

Nvidia also pointed to results published on SemiAnalysis AgentX, a benchmark built around recorded agentic coding sessions rather than single-request tests. It said Vera Rubin NVL72 delivered up to 30 times higher throughput per megawatt than GB300 NVL72 on the DeepSeek V4 Pro model.

Those results also implied up to 45 times lower cost per million tokens, reflecting how benchmark design is changing alongside the shift from simple chatbot queries to more complex agent workflows involving long context, tool calls and sub-agents.

Several software and infrastructure companies have also published benchmark results for the Nvidia Vera CPU. Nvidia cited findings from Perplexity, DeepInfra, Redpanda, Starburst, Kinetica, ClickHouse, Daytona and Prime Intellect across workloads including sandboxing, orchestration, analytics and data streaming.

The results included claims of faster query throughput, lower latency and better performance under parallel workloads, though methodologies differed by company and application.

Scale concerns

As AI factories grow to hundreds of thousands of GPUs, Nvidia said reliability is becoming increasingly tied to economic performance. The company introduced NVLink 6 as a system designed to detect and isolate faults before they affect applications.

The design includes error correction, retry mechanisms, dynamic routing and link rebalancing, intended to prevent local failures from spreading across large AI clusters.

"The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt," Buck said.