
AWS and NVIDIA Plan 2 Million More GPUs: What It Means for AI Infrastructure
PublishedArticle Summary
AWS and NVIDIA plan to deploy 2 million additional NVIDIA GPUs across AWS infrastructure in 2027-2028.
The expanded partnership covers far more than GPUs, including CPUs, networking, AI models, data processing and robotics.
For enterprises, the announcement highlights the growing importance of scalable, integrated AI infrastructure.
“The biggest challenge in AI may no longer be what the models can do. It may be whether there is enough infrastructure to run them.”
That question is becoming increasingly relevant as AI moves from experiments and pilot projects into everyday business operations.
AWS and NVIDIA have announced plans to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure during 2027 and 2028. For Exeton and the enterprises it serves, the announcement is significant because it reflects a broader shift: AI is becoming an infrastructure-intensive business requirement, not simply a software project.
What exactly did AWS and NVIDIA announce?
The headline number is substantial: 2 million additional NVIDIA GPUs planned across AWS infrastructure in 2027-2028.
The expansion is expected to include NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs, while AWS also plans to bring NVIDIA Vera CPUs to its cloud infrastructure.
But the partnership goes well beyond adding more processors.
Infrastructure Layer | AWS and NVIDIA Expansion |
GPUs | Blackwell Ultra, Rubin, Rubin Ultra |
CPUs | NVIDIA Vera |
Networking | NVIDIA Spectrum and AWS EFA |
Memory | NVIDIA NVHBM |
AI Models | NVIDIA Nemotron |
Data Processing | NVIDIA cuDF and cuVS |
Robotics | Jetson, Isaac, and Omniverse |
The companies also plan to build secure AI infrastructure for the U.S. government, including 100,000 GPUs for federal and national-security workloads.
Importantly, these are future deployment plans, not a statement that all 2 million GPUs are already operational.
Why is demand for AI GPUs growing so quickly?
AI workloads are becoming much larger and more varied.
A few years ago, many organizations were still testing generative AI. Today, businesses are looking at AI agents, automated workflows, scientific computing, large-scale inference, robotics and other production workloads.
Running these systems requires much more than a powerful processor.
Think of modern AI infrastructure like an industrial facility. The GPU is important, but it also needs fast memory, networking, storage, power, cooling and software to work effectively.
That is why the AWS-NVIDIA announcement matters. It shows that cloud providers are preparing infrastructure for AI workloads at a scale that would have seemed extraordinary only a few years ago.
Is this just about NVIDIA GPUs?
Not really.
The more interesting part of the announcement is the move toward a full AI infrastructure stack.
NVIDIA Vera CPUs are intended to handle CPU-intensive portions of AI workloads, while high-speed networking helps large numbers of processors communicate efficiently. NVIDIA's NVHBM technology and NVLink Fusion are aimed at improving memory and scale-up capabilities.
On the software side, NVIDIA Nemotron models will be available through AWS services, while CUDA-X libraries such as cuDF and cuVS are being used to accelerate data processing and vector indexing.
The message is clear: simply adding more GPUs isn't enough. The surrounding infrastructure has to keep up.
What does this mean for enterprise AI infrastructure?
For enterprises, the announcement is a reminder that AI infrastructure planning cannot stop at choosing a GPU.
Organizations deploying serious AI workloads may need to consider:
GPU and CPU requirements
Memory capacity and bandwidth
High-speed networking
Storage and data pipelines
Power and cooling
Security requirements
Software compatibility
Future scalability
The right configuration ultimately depends on the workload.
For organizations evaluating enterprise GPUs, the NVIDIA H200 for enterprise AI is one example of how GPU selection can be driven by memory capacity, performance requirements and enterprise use cases rather than specifications alone.
Will every company need thousands of GPUs?
No.
The 2-million-GPU figure describes infrastructure being deployed at cloud scale. Most businesses will have very different requirements.
A professional working with AI development may need a workstation with a powerful GPU. An enterprise running production inference could require dedicated GPU servers, while a large AI research organization may need multi-GPU clusters.
The important question isn't “How many GPUs does everyone need?”
It is:
“What type of infrastructure matches the workload?”
For teams evaluating local AI workstations, Exeton's comparison of NVIDIA GPUs for AI workstations provides useful context on how different GPU classes serve different professional workloads.
What should businesses take away from the AWS-NVIDIA expansion?
There are several practical lessons.
First, AI infrastructure demand is not slowing down. Cloud providers are making enormous capacity commitments because customers are moving AI workloads into production.
Second, GPU selection should be workload-driven. The newest GPU isn't automatically the best choice for every organization.
Third, the supporting infrastructure matters. Networking, memory, storage, power and cooling can become bottlenecks when AI systems scale.
Finally, scalability needs to be considered from the beginning. An infrastructure design that works for a pilot may not work when thousands of users or large production workloads arrive.
Where does Exeton fit into this changing landscape?
As AI infrastructure becomes more complex, enterprises increasingly need to think about the entire deployment rather than an individual component.
That means moving from:
Workload → GPU → server architecture → networking → storage → power/cooling → deployment → scalability
This broader approach is particularly important for organizations building private or hybrid AI infrastructure.
Exeton's work around enterprise AI infrastructure and NVIDIA solutions reflects this wider infrastructure perspective, where hardware decisions need to align with the organization's actual AI requirements.
What does 2 million GPUs mean for the future of AI?
The biggest takeaway isn't simply the number 2 million.
It is what that number says about where AI is heading.
AI workloads are becoming large enough to require massive investments in computing, networking, memory and data-center infrastructure. Cloud providers are preparing for that demand, while enterprises are deciding which workloads should run in the cloud and which may justify dedicated infrastructure.
The AI race is increasingly becoming an infrastructure race.
For businesses, the question isn't simply “Which GPU is fastest?” It is “What infrastructure can support our AI workloads today while giving us room to scale tomorrow?”
As AI becomes more operationally important, those infrastructure decisions will increasingly become business decisions, not just IT decisions. Exeton will continue to be part of that conversation as enterprises evaluate the technologies needed to turn AI ambitions into production-ready infrastructure.