The Pulse
CoreWeave Puts 504 Vera Rubin GPUs Into Seven-Rack Cluster
CoreWeave has placed seven NVIDIA Vera Rubin NVL72 racks into production across two regions, creating a 504-GPU deployment. The company says its software, networking, cooling and storage systems are designed to make the racks operate as one

AI.info Team ·
CoreWeave has brought seven NVIDIA Vera Rubin NVL72 racks into production across two regions, a deployment totaling 504 Rubin GPUs as the cloud provider moves beyond single-rack systems. The figure, reported by TechTarget, puts CoreWeave among the first operators to expose NVIDIA’s newest rack-scale platform as a larger, multi-rack cloud system.
CoreWeave announced the deployment on September 16, 2026, saying customers can run training and inference across hundreds of Rubin GPUs through a single scale-out cluster. The company is targeting agentic AI workloads, where repeated model calls, tool use and intermediate computation can magnify delays in networking, storage and checkpoint handling.
Seven racks, 504 GPUs and one operating domain
Each Vera Rubin NVL72 rack combines 72 NVIDIA Rubin GPUs with 36 Vera CPUs, NVIDIA NVLink 6, ConnectX-9 SuperNICs and BlueField-4 DPUs. CoreWeave connects the racks with NVIDIA Spectrum-X Ethernet networking, extending the high-bandwidth system inside each rack into a larger fabric that can carry traffic between racks.
“CoreWeave was the first AI cloud provider to validate and bring up a Vera Rubin NVL72, demonstrating that this advanced rack-scale architecture could operate as a reliable, high-performance cloud service,” said Chen Goldberg, executive vice president of product and engineering at CoreWeave. “With multi-rack Vera Rubin, we are connecting hundreds of Rubin GPUs as a single scale-out cluster.”
The company says the deployment is intended to support larger model training, demanding inference and reinforcement learning workloads. Its stated goal is to hide the physical boundary between racks from customer jobs, so the workload sees one validated pool of compute rather than a collection of separately managed machines.
Networking becomes part of the computer
CoreWeave’s technical account of the deployment describes a two-tier, non-blocking network fabric with multiple rails and planes. Each Rubin GPU uses two ConnectX-9 SuperNICs, providing up to 1.6 terabits per second of backend connectivity per GPU.
That bandwidth is aimed at synchronized workloads in which GPUs exchange data continuously. CoreWeave says the fabric can support roughly 128,000 GPUs per rail, an architectural capacity rather than the size of the company’s current Rubin installation. The provider also tests traffic over the backend network instead of relying only on faster internal GPU links, using the process to expose faults in switches, cabling, software and other network components.
In its technical account of the bring-up, CoreWeave says it runs synchronized GPU workloads, training-style tests, computational stress tests and custom benchmarks before a rack enters production. The validation process checks complete systems rather than stopping at component-level diagnostics.
Racky and Valvey extend control beyond the servers
CoreWeave has expanded its Mission Control platform with Racky, a rack-management control layer, and Valvey, a programmable liquid-cooling control system. The Rack LifeCycle Controller coordinates hardware detection, firmware updates, metadata, validation, power and cooling as new infrastructure moves from installation toward production.
The approach reflects the physical demands of the Vera Rubin platform. A single NVL72 rack includes high-speed networking, substantial power delivery and liquid cooling designed to operate at 45 degrees Celsius. At multi-rack scale, CoreWeave says a slow GPU, unstable network link, cooling anomaly or configuration mismatch can affect the performance of the wider workload.
Rather than treating cooling as a separate facilities task, CoreWeave puts coolant flow and environmental sensing into the same operational process as compute and networking. That gives operators a way to compare system health across racks while workloads run under sustained load.
Storage changes target the delays between model steps
The Vera Rubin announcement also includes two additions to CoreWeave AI Object Storage: cross-region write acceleration and a lower-cost Archive tier. Cross-region write acceleration lets a job write checkpoints locally while the platform replicates the data to a second region in the background.
CoreWeave says the feature allows a Kubernetes cluster to continue training without waiting for a remote write to complete. Applications can use one bucket across regions without changing code, while permissions and retention rules remain consistent.
The company’s Local Object Transport Accelerator, or LOTA, provides managed caching on CoreWeave Kubernetes Service nodes. CoreWeave says LOTA can deliver reads at local NVMe speeds, reduce latency by eight times compared with direct reads from a traditional storage cluster and provide up to 7 GB per second of throughput per GPU under suitable workload and cache conditions.
“Our datasets span multiple regions, and we can’t afford to have our training schedule dictated by cross-region retrieval delays,” said Cécile Robert-Michon, director of internal infrastructure at Cohere. “CoreWeave AI Object Storage gives us a unified dataset footprint across regions with reads cached locally, so nothing waits on the network.”
The operational test is keeping 504 GPUs moving together
CoreWeave’s announcement frames the deployment as an infrastructure problem as much as a hardware rollout. The provider must keep compute, networking, storage, power, cooling, firmware and software within the same performance envelope while jobs span multiple racks and regions.
That distinction matters because the 504-GPU figure describes production capacity, not simply the number of accelerators installed. The practical measure is whether a distributed training or inference job can continue when a GPU slows, a network path degrades or a rack requires intervention. CoreWeave’s multi-rack Rubin system is now built around that test: seven NVL72 cabinets operating as one managed pool across two regions.