AWS and NVIDIA Plan to Add 2 Million GPUs Between 2027 and 2028

2026-08-28 11:28
Favorite

en.Wedoany.com Reported - Amazon Web Services (AWS) and NVIDIA have announced a major expansion of their strategic collaboration to address the continued growth in demand for global artificial intelligence infrastructure. The two companies plan to deploy an additional 2 million NVIDIA GPUs across AWS's global infrastructure, deepening their collaboration in areas such as AI factories, CPUs, networking, open models, data processing, and robotics, to deliver jointly designed AI solutions to customers and accelerate AI development and deployment.

AI workloads are currently transitioning from pilot phases to production, with applications extending to agentic AI, scientific discovery, enterprise automation, and robotics. Customers need a broader selection of models, faster data processing pipelines, and capabilities for new use cases such as physical AI, while requiring underlying infrastructure that supports the pace of innovation while maintaining the highest levels of security and reliability for mission-critical workloads.

Building on 16 years of joint innovation, AWS and NVIDIA will expand AI computing capabilities and deliver new jointly designed solutions to customers more quickly. AWS offers the broadest range of GPU instances among all cloud providers. At the 2026 NVIDIA GTC conference, AWS announced plans to add more than 1 million NVIDIA GPUs starting in 2026, and actual demand has since exceeded that expectation. AWS now plans to deploy an additional 2 million NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs across AWS's global infrastructure, including AI factories, between 2027 and 2028, covering customer workloads such as agentic AI, scientific discovery, enterprise automation, and physical AI. AWS will also expand NVIDIA Blackwell capacity, including NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances. G7 instances deliver 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 instances, and AWS is the first major cloud provider to offer compute instances accelerated by the RTX PRO 4500. In addition, the two companies are collaborating on NVIDIA Spectrum networking to optimize network performance across GPU clusters for large-scale AI training workloads.

In CPU computing, AWS and NVIDIA are planning to bring infrastructure based on the NVIDIA Vera CPU to AWS, providing additional options for agentic AI workloads that require high-performance CPU compute working in tandem with accelerated infrastructure. Vera is designed for next-generation AI and complements AWS's broad compute selection strategy, spanning from its own custom silicon to partners' latest accelerators and CPUs.

In chip interconnect and memory, AWS announced at re:Invent 2025 support for NVIDIA NVLink Fusion high-speed chip interconnect technology in its next-generation Trainium chips. NVIDIA and Amazon Annapurna Labs are expanding this support by working with memory suppliers to bring NVIDIA's new custom high-bandwidth memory (NVHBM) technology to Trainium, enabling access to faster, more energy-efficient memory. Combined with NVLink Fusion, Annapurna Labs can leverage NVIDIA custom memory technology and scale-up architectures to improve AI workload performance and efficiency, while integrating Trainium and GPUs in a common rack-scale architecture.

For government customers, AWS and NVIDIA plan to build AI factories for the U.S. government, delivering the NVIDIA AI software stack, and plan to provide 100,000 GPUs on secure AWS infrastructure for federal and national security workloads. This collaboration enables government agencies to deploy AI at scale for workloads at Impact Level 6 (IL6) and above.

This expanded collaboration builds on multiple technology integrations already in place. All NVIDIA GPU-based and Trainium-based Amazon EC2 instances, including those leveraging NVLink Fusion, are built on the AWS Nitro System and interconnected via Elastic Fabric Adapters (EFA). GPU-accelerated and Trainium-based EC2 instances will continue to be built on the Nitro System and scaled through EFA. Together, Nitro and EFA ensure that as AWS expands its NVIDIA GPU fleet and integrates new interconnect technologies, customers continue to receive the security, reliability, and network performance that production-grade AI workloads depend on. On the software ecosystem front, the NVIDIA Nemotron family of open models is available on Amazon Bedrock as fully managed, serverless models, and on Amazon SageMaker for customers to deploy and fine-tune on their own.

Data processing acceleration is another collaboration outcome already in place. AWS and NVIDIA are delivering GPU-accelerated data processing on Amazon EMR using Amazon EC2 G7 instances and the NVIDIA cuDF library, delivering up to 3.7 times faster processing and 30% better price-performance compared to CPU-based configurations. In response to the large-scale vector database demands driven by AI applications, retrieval-augmented generation pipelines, and semantic search, GPU-accelerated vector indexing on Amazon OpenSearch Service offloads index construction to dedicated GPUs, delivering up to 9 times faster vector indexing at one-quarter of the cost, available for both managed clusters and Amazon OpenSearch Serverless.

In robotics, Amazon Robotics is collaborating with NVIDIA to integrate NVIDIA's full-stack physical AI platform for developing next-generation robots. The platform includes NVIDIA Jetson, NVIDIA Omniverse libraries, and the NVIDIA Isaac open robotics development platform. The collaboration covers simulation, synthetic data generation, robot training, route optimization, functional safety, and real-to-simulation validation, all running on GPU-accelerated Amazon EC2 instances to support the large-scale simulation, diverse training data, and continuous real-world validation required for robotics workloads.

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com