Nvidia's Vera Rubin Chip Aims to Control Every Layer of AI Data Centers
Nvidia is making a strategic shift from GPU-only supplier to a complete AI infrastructure provider, unveiling its Vera Rubin chip system that pairs CPUs with GPUs to orchestrate complex AI operations. The company revealed new performance benchmarks this week, showing that Vera Rubin processes 10 times as many tokens per watt as its previous-generation Grace Blackwell system, positioning the chipmaker to capture more of the lucrative AI data center market.
Why Is Nvidia Moving Into CPUs Now?
For years, Nvidia dominated AI infrastructure by making graphics processing units, or GPUs, which are the workhorses for training and running large language models. But as AI systems have become more sophisticated, especially with the rise of AI agents that need to coordinate multiple tasks, the industry has increasingly demanded CPUs, or central processing units, to handle orchestration, networking, and data flow management.
This shift explains why Nvidia executives have been emphasizing their role as a supplier of complete AI systems rather than just specialized chips. Ian Buck, Nvidia's vice president of accelerated computing and architect of the company's CUDA software platform, explained the company's philosophy during a technical workshop at Nvidia's Santa Clara headquarters.
"We're on a road map to crank out new architectures, not just GPUs but CPUs. We're going to keep innovating, because it's do this or die in Silicon Valley," said Ian Buck.
Ian Buck, Vice President of Accelerated Computing at Nvidia
The Vera Rubin system is designed with one CPU for every two GPUs. In a single Vera Rubin NVL72 super chip system, there are 36 Vera CPUs paired with 72 Rubin GPUs, creating a tightly integrated platform for AI workloads.
What Makes Vera Rubin Different From Previous Nvidia Systems?
Nvidia has engineered several technical improvements into Vera Rubin that address pain points data center operators have experienced with earlier systems. The company is marketing the new platform as "cable-free compute" and "hot-swappable," meaning customers can theoretically reduce installation time from a couple of hours to just a few minutes.
Beyond installation convenience, Vera Rubin offers substantial performance gains. The system features localized memory subsystems that provide nearly three times as much memory bandwidth as Blackwell, addressing an ongoing shortage of high-bandwidth memory that has constrained AI deployments. The entire system is 100 percent liquid-cooled, which reduces energy consumption compared to air-cooling methods.
A key architectural decision sets Vera Rubin apart from competitors like AMD. Rather than using a chiplet architecture, where multiple smaller chips are stitched together, Nvidia built Vera Rubin as a single, monolithic chip. According to the company, this design choice eliminates what it calls "a heavy tax on memory bandwidth and data movement," allowing data to flow more quickly across the integrated circuit.
How to Evaluate Vera Rubin's Competitive Position
- Performance Metrics: Vera Rubin processes 10 times as many tokens per watt as Grace Blackwell, and Nvidia claims its Vera CPU outperforms rival CPUs from AMD and Intel at agentic AI tasks, though the company used slightly older generations of competitors' processors in its benchmarks.
- Installation and Operations: The system reduces cable requirements significantly and supports hot-swappable components, potentially cutting deployment time from hours to minutes and lowering operational complexity for large-scale data centers.
- Memory Bandwidth: Vera Rubin delivers nearly three times the memory bandwidth of Blackwell, a critical advantage as AI labs and cloud providers struggle with high-bandwidth memory constraints that limit model training and inference speeds.
- Thermal Efficiency: Full liquid cooling reduces energy consumption compared to traditional air-cooled systems, lowering operational costs for data centers running continuous AI workloads.
Nvidia is particularly sensitive to any perception of delays after its Blackwell chips reportedly overheated when connected in the company's customized server racks, forcing design changes and pushing back shipments. CEO Jensen Huang has repeatedly stated that Vera Rubin is ramping to full production and will ship in the second half of 2026, with early customers including Microsoft, OpenAI, and Oracle.
The timing of Nvidia's Vera Rubin push is strategic. The company is unveiling these benchmarks just ahead of rival AMD's annual conference, where AMD executives are expected to showcase their next-generation Helios AI chip rack designed to compete directly with Nvidia's offerings. Both companies are vying for large-scale, multiyear contracts to supply chips to AI hyperscalers like Meta and Amazon, as well as AI labs including OpenAI, Anthropic, and SpaceX AI.
During a tour of Nvidia's Silicon Valley data center lab, executives revealed that OpenAI already has one Vera Rubin rack in use, signaling confidence from one of the world's most demanding AI customers. Nvidia is also selling the Vera CPU as a standalone product and has told Chinese customers that these processors could be ready as soon as August.
The shift toward hybrid CPU-GPU systems reflects a broader industry evolution. As AI moves beyond simple model training toward more complex agentic systems that coordinate multiple operations, the demand for CPUs that can manage orchestration and data flows has grown substantially. By positioning itself as a supplier of complete systems rather than individual components, Nvidia is attempting to deepen its lock on the AI infrastructure market and maintain its dominance even as competitors like AMD gain ground in traditional data center CPUs.