AWS and NVIDIA Are Deploying 2 Million More GPUs to Power the Next Wave of AI Robots and Agents
AWS and NVIDIA announced a major expansion of their 16-year partnership, committing to deploy 2 million additional GPUs across AWS's global infrastructure by 2028 to meet surging demand for AI infrastructure. The collaboration will deepen across multiple layers, including graphics processing units (GPUs), central processing units (CPUs), networking, open-source models, data processing, and robotics, enabling customers to scale AI workloads from pilot projects into full production.
Why Are AWS and NVIDIA Expanding Their Partnership Now?
The explosion in AI adoption has outpaced even the most optimistic forecasts. Customers are moving rapidly from testing AI systems to deploying them at scale for real-world applications like autonomous agents, scientific discovery, enterprise automation, and physical AI (robots and autonomous systems). AWS had already announced plans to add more than 1 million NVIDIA GPUs starting in 2026, but demand has exceeded those expectations, prompting the additional 2 million GPU commitment for 2027-2028.
"Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together," said Matt Garman, CEO of AWS.
Matt Garman, CEO of AWS
Jensen Huang, founder and CEO of NVIDIA, emphasized the scale of the opportunity. "NVIDIA and AWS have built one of the great growth engines of the AI era, and demand is running ahead of every forecast," he stated. "For 16 years, we have scaled NVIDIA computing in the cloud together. Now, we are expanding our partnership across the full stack to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver".
What Specific Infrastructure Improvements Are Planned?
The expanded collaboration spans multiple technical domains designed to optimize AI workloads across different use cases. AWS will deploy NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs, which represent the latest generation of NVIDIA's accelerator hardware. The company will also expand capacity for NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances, which deliver 4.6 times the AI inference performance and 2.1 times the graphics performance compared to previous-generation G6 instances.
- GPU Deployment: 2 million additional NVIDIA GPUs across AWS global infrastructure in 2027-2028, including Blackwell Ultra, Rubin, and Rubin Ultra models for diverse AI workloads
- CPU Integration: NVIDIA Vera CPU-based infrastructure coming to AWS to support agentic AI workloads requiring high-performance CPU compute alongside GPU acceleration
- Memory Technology: NVIDIA NVLink Fusion with custom high-bandwidth memory (NVHBM) integrated with AWS Trainium chips for faster, more power-efficient AI processing
- Networking Optimization: NVIDIA Spectrum networking collaboration to enhance network performance for large-scale AI training across GPU clusters
- Government AI Infrastructure: 100,000 GPUs on secure AWS infrastructure for federal and national-security workloads classified at Impact Level 6 and above
AWS is also the first major cloud provider to offer compute instances accelerated by RTX PRO 4500 Blackwell Server Edition GPUs, giving customers access to cutting-edge graphics and AI capabilities.
How to Leverage These Infrastructure Improvements for Your AI Projects
- Evaluate Your Workload Type: Determine whether your AI application requires GPU acceleration for training large models, CPU-intensive agentic AI tasks, or specialized graphics processing, then select the appropriate AWS instance type from the expanded portfolio
- Plan for Scale: If you are currently running pilot AI projects, begin planning your production deployment strategy now, as the 2 million GPU expansion will increase availability and reduce potential capacity constraints through 2028
- Explore Open Models: Take advantage of NVIDIA Nemotron open models available on Amazon Bedrock and Amazon SageMaker, which offer more model choice and flexibility compared to proprietary alternatives
- Optimize Data Pipelines: Integrate NVIDIA cuDF and cuVS CUDA-X libraries with Amazon EMR and Amazon OpenSearch to accelerate data processing and vector indexing for faster, more cost-efficient analytics
- Consider Robotics Applications: If your organization is exploring warehouse automation or next-generation robotics, AWS and NVIDIA's expanded physical AI platform support through Amazon Robotics provides a production-ready foundation
What Does This Mean for Enterprise AI Adoption?
The partnership expansion signals that AWS and NVIDIA expect agentic AI and physical AI to become mainstream enterprise technologies within the next two to three years. Agentic AI refers to autonomous systems that can plan, reason, and take actions with minimal human intervention. Physical AI extends this concept to robots and autonomous machines operating in the real world. By committing to 2 million additional GPUs, AWS and NVIDIA are betting that demand for these technologies will far exceed current capacity.
The collaboration also addresses a critical pain point for enterprises: infrastructure reliability and security. All NVIDIA GPU-based and AWS Trainium-based EC2 instances will be built on the AWS Nitro System and interconnected through the Elastic Fabric Adapter (EFA), which together ensure security, reliability, and network performance for production AI workloads at scale. This is particularly important for government agencies, which will have access to 100,000 GPUs on secure AWS infrastructure for classified workloads.
The expansion of NVIDIA Nemotron open models on Amazon Bedrock and Amazon SageMaker also reflects a broader industry trend toward open-source AI. Customers can now choose between fully managed, serverless models on Bedrock or deploy and fine-tune models on their own infrastructure using SageMaker, giving them flexibility in how they build and operate AI systems.
For data-intensive applications, the integration of NVIDIA cuDF and cuVS CUDA-X libraries with Amazon EMR and Amazon OpenSearch will accelerate data processing and vector indexing, enabling faster and more cost-efficient analytics and AI applications. This is particularly valuable for organizations building large-scale recommendation systems, search engines, or other applications that rely on vector similarity search.
The partnership's focus on robotics through Amazon Robotics' adoption of NVIDIA's physical AI platform signals that warehouse automation and next-generation robots are moving from research projects into production deployments. This expansion will likely accelerate innovation in logistics, manufacturing, and other industries where physical automation can deliver significant productivity gains.