IT Brief UK - Technology news for CIOs & IT decision-makers
United Kingdom
NVIDIA ramps up Vera Rubin AI system for cloud giants

NVIDIA ramps up Vera Rubin AI system for cloud giants

Wed, 22nd Jul 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Production of NVIDIA's Vera Rubin NVL72 AI system is ramping up, with racks already running at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure.

The rollout marks the commercial arrival of NVIDIA's latest rack-scale AI platform, backed by a supply chain spanning more than 350 factory sites across 30 countries.

Vera Rubin NVL72 is the centrepiece of a broader system design that combines seven chips and five rack trays into a single package. NVIDIA said the platform was engineered as an integrated system rather than assembled from standard components, spanning compute, networking and cooling.

At the processor level, the setup includes the NVIDIA Vera CPU at the centre of the platform. NVIDIA said the chip was designed for AI agent workloads and claimed gains in single-threaded performance, core-to-core bandwidth and memory latency over competing chiplet designs.

Networking is also central to the launch. NVIDIA said its sixth-generation NVLink scale-up network delivers more than twice the throughput on complex workloads, along with lower latency and higher packet rates than standard Ethernet. Its Spectrum-X Ethernet offering is designed to link systems across larger clusters and multiple sites.

NVIDIA also said several infrastructure groups, including CoreWeave, Microsoft, SpaceXAI and Tesla, are among the first to deploy Spectrum-6 switches for AI workloads. CoreWeave, Lambda and Oracle Cloud Infrastructure are among early adopters of NVIDIA's photonics-based networking hardware, the company said.

Benchmark claims

One of the first public performance results came from CoreWeave, which said it had brought up and validated Vera Rubin NVL72 before publishing benchmark data from live hardware. In a DeepSeek-R1 benchmark, CoreWeave said the system delivered 10 times more tokens per second per megawatt than Grace Blackwell NVL72.

CoreWeave said the test focused on power efficiency, a key constraint for data centre operators building large AI clusters. NVIDIA also said Google Cloud's first A5X instance based on Vera Rubin NVL72 is running for London startup Ineffable Intelligence.

According to NVIDIA, Google Cloud's A5X system uses Vera Rubin NVL72 with Google Virgo networking for data centre scale-out. The bare-metal instances are intended to support reinforcement learning workloads and can scale from large single-site clusters to multi-site deployments, NVIDIA said.

"The next era of research requires the next era of hardware," said Lasse Espeholt, co-founder of Ineffable Intelligence. "We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners."

Europe focus

NVIDIA also tied Vera Rubin to a broader European infrastructure push through Microsoft and Mistral. The platform will underpin a newly expanded partnership between the two companies, centred on a multibillion-dollar agreement to expand AI infrastructure in Europe, NVIDIA said.

Under the arrangement, Mistral is adding GPU capacity based on thousands of Vera Rubin processors to expand AI compute availability for customers. NVIDIA said the system will support Mistral Compute and Microsoft's European AI infrastructure, with the broader buildout drawing on tens of thousands of GPUs.

The move reflects growing demand in Europe for AI services that can operate within regional rules on data control, governance and operational independence. NVIDIA said Vera Rubin's architecture is intended to support open-model deployments across public cloud, cloud-connected and fully disconnected private cloud environments.

Mistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, according to NVIDIA, and Mistral models have been integrated into Microsoft Copilot Studio. Azure Local and Foundry Local allow customers to use the same models and tools across cloud and customer-controlled environments, NVIDIA said.

CPU and cooling

NVIDIA also highlighted benchmark data from DeepInfra focused on the Vera CPU rather than the full rack system. According to NVIDIA, DeepInfra found the processor more than twice as fast in orchestration tasks and able to support up to 1.6 times more concurrent AI agents at the same quality of service than alternative CPUs.

DeepInfra processes nearly five trillion tokens a week, NVIDIA said, with about 30% of that driven by agentic systems. The benchmark was designed and run by DeepInfra using its production AI agent infrastructure, NVIDIA added.

On the physical design side, NVIDIA said three generations of rack-scale co-design had removed cables, fans and hoses from the compute tray, cutting assembly time from hours to one minute. It also said a liquid-cooling inlet temperature of 45 degrees Celsius allows dry-cooler operation without chillers, which can save millions of gallons of water per megawatt each year in new AI factories.

NVIDIA said the broader objective is to address the economics of increasingly complex AI systems, particularly agentic workloads that require more compute and tighter control over power use. CoreWeave's benchmark result, highlighted in the launch, framed that issue in practical terms by measuring tokens produced for each megawatt of electricity consumed.