India's Yotta Bets $12 Billion on 80,000 Nvidia Vera Rubin GPUs, Signaling Shift in AI Infrastructure Race
India's cloud operator Yotta Data Services just placed one of the largest GPU orders ever disclosed publicly, committing $12 billion to deploy 80,000 Nvidia Vera Rubin graphics processors across two new data centers. The announcement, made on September 11, 2026, at the GPU Future Forum in Mumbai, marks a turning point in how AI infrastructure is being built outside the United States, and it reveals what Nvidia's next generation of chips is actually designed to do.
Vera Rubin is Nvidia's successor to the current Blackwell generation, and it represents a fundamental shift in how the company is thinking about AI hardware. Rather than simply making chips faster at raw computation, Nvidia focused on memory bandwidth, the speed at which data moves between storage and processors. This matters because modern artificial intelligence models spend as much time waiting for data as they do computing on it.
What Makes Vera Rubin Different From Previous Nvidia Chips?
Each Vera Rubin GPU packs 288 gigabytes of HBM4 memory, delivering roughly 22 terabytes per second of bandwidth per chip. That is about 2.75 times faster than the roughly 8 terabytes per second bandwidth of Blackwell Ultra's HBM3e memory, even though both generations hold the same total memory capacity. For AI workloads, this bandwidth leap often matters more than raw processing power on a spec sheet.
The GPU itself is built on TSMC's 3-nanometer process as a dual-die package with roughly 336 billion transistors, rated at approximately 50 petaflops of NVFP4 inference performance per chip. It pairs with a new Vera CPU, an 88-core Arm-based processor design with simultaneous multithreading that gives 176 threads per socket, connected to the GPU over NVLink-C2C at roughly 1.8 terabytes per second.
Nvidia entered full production on Vera Rubin around June 1, 2026, with partner shipments scheduled to begin in the second half of the year, which lines up with Yotta's own rollout timeline.
How Does Vera Rubin Perform on Real-World AI Agent Workloads?
The real-world performance advantage of Vera Rubin became clear when independent semiconductor research firm SemiAnalysis released benchmark results on September 15, 2026. The AgentX benchmark, which tests how well hardware handles AI agents performing complex tasks like research and analysis, showed Vera Rubin NVL72 delivering up to 30 times higher throughput per megawatt than Nvidia's own GB300 NVL72 on agentic-coding inference workloads.
This benchmark matters because it reflects how AI is actually being used in production. An agentic AI session looks nothing like a single chat turn. An agent researching a company for an investment decision queries financial databases, searches filings, invokes sub-agents to run peer comparisons, synthesizes multiple sources, and loops through tool calls, all before producing a final answer. Across 100 trillion tokens of real-world traffic analyzed by OpenRouter, single agentic requests consume roughly 15 times the tokens of an ordinary chat interaction.
The benchmark corpus consisted of 393 anonymized real-world coding-agent traces, preserving the real prompt sizes, tool-call timing, sub-agent fan-out, and inter-turn delays of production agentic traffic. The median session in the corpus sends 142,000 input tokens and 444 output tokens per request, and 44 percent of sessions spawn sub-agents.
Why Power Efficiency Has Become More Important Than Raw Speed?
For most of AI's recent commercial history, the dominant procurement metric was peak theoretical FLOPS (floating-point operations per second) per GPU, followed by tokens per second per chip. Those metrics were appropriate when compute availability was the binding constraint. Increasingly, power is.
New power infrastructure takes years to permit and build. Cooling capacity limits are forcing data center operators to manage thermals more aggressively. In this environment, the question governing AI factory revenue is not how many FLOPS a GPU delivers in a controlled test, but how many tokens it produces per megawatt of available grid power, continuously, at scale, under realistic workloads.
"The metric for AI infrastructure is fast shifting from peak performance to validated agentic tokens per megawatt," said Ian Buck, VP of Hyperscale and HPC at Nvidia.
Ian Buck, VP of Hyperscale and HPC at Nvidia
On the DeepSeek V4 Pro model, a 1.6-trillion-parameter Mixture-of-Experts architecture, Nvidia's results show Vera Rubin NVL72 delivering 30 times higher throughput per megawatt than GB300 NVL72 under AgentX conditions. The cost advantage is steeper still: token costs are 45 times lower per million.
How Is Yotta Structuring Its $12 Billion Investment?
Yotta's plan splits 80,000 GPUs across two sites in India, a geographic and technical hedge that spreads the power and cooling burden across two grid regions. This matters given how much electricity 80,000 GPUs actually consume once racked and running.
- Greater Noida Site: A new D4 data center under construction will house 40,000 Vera Rubin GPUs inside a 120-megawatt facility, representing Nvidia's newest generation of chips.
- Navi Mumbai Site: An NM2 campus expansion will run 40,000 Nvidia GB300 Blackwell Ultra GPUs across 80 megawatts, building on the existing Blackwell Ultra generation.
- Combined Capacity: The two sites add 200 megawatts of GPU-dense capacity, putting Yotta in a similar range to mid-sized hyperscale campuses being built in the US and Middle East.
The $12 billion figure covers land, power infrastructure, liquid cooling, networking, and the GPUs themselves, though Yotta has not broken out how much goes to Nvidia hardware specifically. CEO Sunil Gupta told Moneycontrol that the Noida site is dedicated to Nvidia's newest Vera Rubin GPUs while Navi Mumbai keeps building on the existing Blackwell Ultra generation, giving Yotta a mixed fleet rather than a single-generation bet.
For India, where dedicated AI data-center capacity barely existed three years ago, a 200-megawatt combined GPU footprint is a genuinely large jump. Gupta says the company already controls a majority share of India's installed GPU capacity, and this expansion is designed to lock that lead in before rivals catch up.
How to Understand the Vera Rubin NVL72 Rack Architecture
Nvidia sells Vera Rubin primarily as a rack-scale system called Vera Rubin NVL72, pairing 72 Rubin GPUs with 36 Vera CPUs in a single liquid-cooled cabinet. Understanding this architecture helps explain why the efficiency gains are so significant.
- Memory Capacity: The rack pools 20.7 terabytes of HBM4 memory, a massive amount of fast storage that allows the system to hold enormous AI models and their associated data in memory simultaneously.
- Internal Bandwidth: Data moves internally at roughly 260 terabytes per second over NVLink 6, allowing the 72 GPUs to share information at speeds that would be impossible over conventional networking.
- Compute Performance: The rack reaches close to 3.6 exaflops of NVFP4 inference throughput, enough to run trillion-parameter models at production scale.
- Power Requirements: Power draw for a full rack lands between 190 and 230 kilowatts, well above the roughly 140 kilowatts a GB300 NVL72 rack draws, and it requires full liquid cooling with no air-cooled configuration on offer.
Yotta's decision to run Vera Rubin at one site and stick with GB300 Blackwell Ultra at the other gives a useful side-by-side comparison. Rack-level memory capacity stays flat between generations at 20.7 terabytes pooled HBM, but bandwidth, compute, and power all jump with Vera Rubin. Rubin roughly triples FP4 compute per GPU and nearly triples memory bandwidth, but it also draws considerably more power per rack. That trade-off is exactly why Yotta is running both generations in parallel rather than betting the entire buildout on the newest chip.
GB300 racks are proven, available now, and cheaper to power. Vera Rubin racks are faster but demand data-center power density India's grid has not had to deliver at this scale before. This mixed approach allows Yotta to balance proven technology with cutting-edge performance as it scales up its AI infrastructure footprint across the country.