Amazon Web Services (AWS) and NVIDIA announced on August 26, 2026, the expansion of their technology partnership, with a plan to deploy an additional two million NVIDIA graphics processing units across AWS’s global infrastructure during 2027 and 2028. The move extends to processors, networking, open models, data processing, and robotics, rather than limiting the collaboration to providing cloud GPUs.
The plan follows AWS’s announcement at NVIDIA GTC 2026 of its intention to add more than one million GPUs starting in 2026, as the companies said demand had exceeded previous expectations. The new capacity will be used for workloads including agentic AI, scientific discovery, enterprise automation, physical AI, and robotics.
Broader Hardware Options for AWS Customers
AWS aims to keep computing options open for customers by combining NVIDIA units with AWS-designed Trainium units, or by using both types together. The expansion includes work to introduce an architecture based on NVIDIA Vera processors to AWS, providing high-performance CPU capacity for agentic AI and reinforcement learning workloads.
Vera is intended for tasks such as code execution, tool use, running sandboxed environments, analytics, data pipelines, and coordination among system components. The companies said the processor can be used as a host unit for accelerated systems or as a standalone processor within AI factories.
Linking Trainium with NVIDIA Technologies
The collaboration also includes support for NVLink Fusion technology for high-speed chip-to-chip connectivity in future Trainium generations. Amazon’s Annapurna Labs will work with NVIDIA on custom high-bandwidth memory technology, NVHBM, which is expected to give Trainium access to faster and more power-efficient memory, in cooperation with memory suppliers.
The combination of NVHBM and NVLink Fusion aims to improve scalability and performance in AI workloads, integrating Trainium and NVIDIA units within a shared rack-level architecture. The companies are also collaborating on networking technologies that connect GPUs more efficiently during large-scale model training.
Infrastructure Designed for the Government Sector
AWS and NVIDIA plan to build AI factories for the U.S. government, including 100,000 GPUs on AWS’s secure architecture, to support national security workloads and classified government missions at Impact Level 6 (IL6) and above. The announcement does not specify a separate timeline for this part or operational distribution details for the factories.
What Changes in Practice?
The move shows that competition in AI cloud infrastructure is not only about the number of GPUs, but also about integrating processors, memory, networking, software, and security levels. Customers benefit from multiple hardware options and deeper integration between NVIDIA and AWS components, but the announced capacity represents a future commitment rather than capacity immediately available.
AWS notes that existing integrations currently provide GPU- and Trainium-based EC2 instances within the AWS Nitro System, connectivity through Elastic Fabric Adapter, as well as Nemotron models on Amazon Bedrock and Amazon SageMaker. It also says that GPU-accelerated data processing on Amazon EMR achieved up to 3.7 times the Apache Spark processing speed and a 30% improvement in performance per cost compared with the alternative configurations, while the accelerated vector index on Amazon OpenSearch Service provided indexing up to 9 times faster at one-quarter the cost of the comparison configurations. Amazon Robotics uses Jetson, Omniverse, and Isaac platforms to develop warehouse automation.