Logo
FrontierNews.ai

Why Big Tech Is Betting on Arm CPUs as the Secret Weapon for AI Agents

Arm-based processors are becoming the backbone of hyperscale AI infrastructure, with Microsoft and Amazon deploying custom chips at unprecedented speed to handle the complexity of AI agents rather than just model training. In back-to-back earnings calls, both companies revealed that Arm architecture is no longer an experiment but a core strategy for managing the computational demands of next-generation artificial intelligence systems.

Why Are Hyperscalers Suddenly Obsessed with CPUs?

For years, the AI conversation centered on graphics processing units (GPUs), which excel at the mathematical operations required to train large language models and generate text. But the shift toward AI agents, systems that can retrieve data, call external tools, run code, and orchestrate multi-step workflows, has changed the equation entirely. These agents are messier and more complex than simple model inference.

Microsoft CEO Satya Nadella captured the architectural shift in nine words: "When it comes to running agents, CPUs are just as important as GPUs." CPUs handle everything surrounding the GPU workloads: routing requests, preparing data, orchestrating tools, managing memory and storage, and keeping thousands of concurrent interactions running smoothly without system failures.

Satya Nadella

The numbers back up the strategic pivot. Amazon reported its chips business, which includes Graviton processors, Trainium accelerators, and Nitro system controllers, has grown to a $25 billion annual revenue run rate, with year-over-year triple-digit growth. More strikingly, Graviton revenue commitments nearly tripled quarter over quarter, and the latest Graviton5 generation is being adopted almost twice as fast as its predecessor, Graviton4.

How Are Microsoft and Amazon Deploying These Custom Chips?

Microsoft's Cobalt 200, built on Arm Neoverse Compute Subsystems V3 architecture, represents the company's second-generation Arm-based cloud processor. The chip delivers up to 50% better performance than its predecessor and scales to 128 virtual CPUs (vCPUs), designed specifically for cloud-native, data-intensive, and agentic workloads. The speed of deployment is remarkable: Microsoft announced Cobalt 200 in November 2025, opened a virtual machine preview in June 2026, and by late July 2026 had racks running in more than 25 data centers worldwide.

AWS is following a similar trajectory with Graviton5. The latest generation delivers up to 25% better compute performance than Graviton4 and up to 35% faster machine learning inference. Graviton5-powered M9g and M9gd instances became generally available in June 2026. Across the entire Graviton lineup, AWS claims 30 to 40% better price-performance compared to equivalent instances from competitors.

The adoption metrics reveal how deeply these custom CPUs have penetrated hyperscale operations. According to Amazon, 98% of its top 1,000 EC2 customers use Graviton. Over 120,000 customers build applications on the platform. For three consecutive years, more than half of all new AWS CPU capacity has been Graviton-based.

What Makes Arm Architecture the Right Choice for AI Infrastructure?

Arm architecture allows hyperscalers to optimize silicon for their own specific workloads and economics while maintaining a shared software foundation underneath. This approach accelerates time-to-revenue and reduces the complexity of managing multiple incompatible processor architectures across massive data center fleets.

The efficiency gains translate directly to business impact. Microsoft drew an explicit line from infrastructure efficiency to financial performance. Azure revenue grew 43% in the recent quarter, with demand still outrunning available capacity. Chief Financial Officer Amy Hood credited efficiency gains across both CPU and GPU fleets, combined with faster infrastructure deployment, for capacity that was monetized almost as quickly as it came online.

At hyperscale, CPU efficiency is not merely a cost-saving measure. When demand exceeds supply, a more efficient fleet squeezes more usable capacity from the same physical racks and the same power envelope. That additional capacity directly becomes revenue. The electricity demands of AI data centers jumped 50% in 2025 alone, according to the International Energy Agency, making every watt of efficiency increasingly valuable.

How to Understand the Broader Shift in AI Infrastructure

  • From Training to Orchestration: The first wave of generative AI focused on model size, accelerator availability, and token generation speed. Agentic AI requires orchestration, data retrieval, tool integration, and policy enforcement, shifting computational demands from GPUs alone to a balanced CPU-GPU architecture.
  • Custom Silicon as Competitive Advantage: Microsoft, Amazon, Google, and NVIDIA have all chosen Arm-based CPUs for their custom silicon strategies, including Google's Axion processors and NVIDIA's Vera CPU, indicating that Arm architecture has become the industry standard for AI infrastructure orchestration.
  • Efficiency as a Capacity Lever: In a supply-constrained environment, CPU efficiency directly unlocks additional computing capacity without requiring new physical infrastructure, making performance-per-watt a critical business metric alongside raw performance.

Google Cloud built Axion, its custom Arm-based CPU family, to serve as the general-purpose compute foundation for its cloud infrastructure and as the AI head node CPU alongside its latest TPU (Tensor Processing Unit) generations. Axion provides the orchestration, networking, and infrastructure services that keep AI systems running at hyperscale.

NVIDIA made the same architectural choice with Vera, its latest Arm-based CPU, which serves as the compute foundation for next-generation AI systems including Rubin. CPUs coordinate memory, networking, and accelerators across rack-scale deployments.

The convergence is striking. Microsoft Cobalt, AWS Graviton, Google Axion, and NVIDIA Vera all point to the same conclusion: as AI infrastructure scales from model inference to autonomous agent execution, Arm-based CPUs are becoming the common compute foundation that orchestrates modern AI data centers.

The market opportunity is substantial. International Data Corporation (IDC) expects global AI infrastructure spending to reach $497 billion in 2026, up roughly 56% year over year. Real demand is now showing up beyond GPU systems, including orchestration, data pipelines, and CPU-only inference clusters. This expansion validates the strategic bet that hyperscalers are making on custom Arm-based processors as the backbone of next-generation AI infrastructure.