AWS and NVIDIA on August 27, 2026 announced the biggest expansion of their cloud-AI partnership to date: 2 million additional NVIDIA GPUs deployed across AWS’s global infrastructure between 2027 and 2028. The news landed less than 48 hours after AWS confirmed a separate deal to acquire DuckLabs, the Amsterdam-based team behind the open-source analytics engine DuckDB. Together, the two moves mark the clearest signal yet that Amazon is racing to lock down both the compute and the data layer of the AI stack before rivals Microsoft and Google pull further ahead.
For engineering teams running workloads on AWS, the announcements matter beyond the headline numbers. GPU supply has been the binding constraint on AI infrastructure for three straight years, and a 2-million-unit commitment reshapes capacity planning, pricing, and regional availability for anyone budgeting Bedrock, SageMaker, or EC2 P-series spend into 2027. The DuckDB acquisition, meanwhile, signals AWS is done treating fast local analytics as a third-party dependency and wants it baked into the data layer that feeds agentic AI systems.
What AWS and NVIDIA Actually Announced on August 27
The core commitment is straightforward on paper: AWS will deploy 2 million additional NVIDIA GPUs across its global data center footprint in 2027 and 2028. That builds directly on top of a prior pledge, made earlier in 2026, to bring more than 1 million NVIDIA GPUs online, including Blackwell and the newer Rubin architecture, across AWS regions. Combined, the two commitments push AWS’s disclosed NVIDIA GPU pipeline toward roughly 3 million accelerators by the end of the decade, according to AWS’s own blog post on the deepened collaboration.
The expansion is not limited to raw chip count. AWS said it will bring NVIDIA’s Vera CPU-based infrastructure into its data centers and extend NVLink Fusion with custom high-bandwidth memory, both aimed at cutting the interconnect bottlenecks that slow down training runs at frontier scale. A separate, security-flavored piece of the deal covers AI factories built specifically for U.S. government customers, including clusters running up to 100,000 GPUs on infrastructure secured through AWS Nitro System and Elastic Fabric Adapter. That detail matters for federal agencies and defense contractors that have been asking cloud vendors for isolated, classified-capable AI capacity rather than shared commercial regions.
NVIDIA framed the deal as a full-stack expansion rather than a chip order. “NVIDIA and @awscloud are expanding the partnership across the full stack — GPUs, CPUs, networking, open models and software — to make agentic and physical AI real at a pace and scale that only we can deliver together,” the company said in a public post about the announcement (source). AWS echoed the scale of the commitment directly: “AWS will deploy 2 million additional NVIDIA GPUs across its global infrastructure in 2027–2028 — and that’s just the start” (source).
The DuckDB Acquisition: Why AWS Wants an Analytics Engine It Doesn’t Already Own
Two days before the GPU news broke, Amazon signed a definitive agreement to acquire DuckLabs, the company behind DuckDB, an open-source, in-process analytical database that has become a favorite among data engineers for sub-terabyte analytical queries. AWS’s own blog post described the deal plainly: “Today we are announcing that Amazon has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind the open-source analytical database DuckDB” (source).
DuckDB’s appeal is speed and simplicity. It runs embedded, in-process, without the operational overhead of a distributed data warehouse, and it has built a large, loyal open-source community since its 2019 debut out of CWI, the Dutch national research institute for math and computer science. AWS’s stated rationale is that DuckDB is now a common runtime for data engineering, data science, and increasingly for AI agents that need to query structured data on the fly rather than wait on a warehouse job. Folding that engine into AWS’s own stack gives Amazon a fast local-query layer to sit next to Redshift and Athena, and positions it as connective tissue for agentic workloads running on Bedrock.
The timing is not incidental. Snowflake and Databricks have both been racing to add native support for smaller, embedded query engines to compete for the same fast-local-analytics use case. By buying the team outright rather than partnering, AWS avoids a licensing dependency and can integrate DuckDB’s execution engine directly into services like SageMaker and Bedrock Knowledge Bases, where response latency for agent-driven queries is a competitive differentiator.
Custom Silicon: Graviton5 Lands as AWS’s AI-Era Chip
The GPU and database news arrived alongside a quieter but consequential milestone: Graviton5, AWS’s fifth-generation Arm-based custom processor, reached general availability in mid-August 2026. AWS describes Graviton5 as built for the agentic AI era, a positioning shift from earlier Graviton generations that were pitched mainly on price-performance for general compute. The company says its custom silicon business, which spans Graviton, Trainium, and Inferentia, has now crossed a $25 billion annual revenue run rate.
That figure is worth sitting with. Custom silicon was, until recently, treated as a cost-control side project for hyperscalers. A $25 billion run rate puts AWS’s chip business in the same conversation as mid-sized independent semiconductor vendors, and it explains why Amazon is comfortable committing to millions of NVIDIA GPUs while simultaneously pushing its own Trainium roadmap. The two aren’t in conflict: Trainium chips absorb price-sensitive training and inference workloads, while NVIDIA GPUs remain the default for customers who need CUDA compatibility or the largest frontier training runs.
Bedrock’s Model Lineup Gets a Frontier Refresh
Alongside the infrastructure news, AWS used its August updates to refresh what’s actually available to run on that infrastructure. Anthropic’s Claude Opus 5 is now live on Amazon Bedrock and the Claude Platform, giving AWS customers direct access to Anthropic’s top-tier model for agentic coding and complex knowledge work. Anthropic’s Claude Fable 5 has also returned to Bedrock, positioned for strong coding, knowledge-work, and vision performance. Separately, AWS has made OpenAI’s Daybreak Red and Daybreak Blue models available to eligible Bedrock customers, specialized models built for vulnerability discovery, exploit reproduction, and mitigation development, aimed squarely at security teams rather than general-purpose chat.
That combination, frontier coding models, security-specialized models, and a fresh GPU pipeline landing in the same month, is not a coincidence. AWS is betting that enterprise AI spend is shifting from experimentation to production deployment, and it wants the compute, the models, and the data engine all available inside one procurement relationship.
Who’s Actually Buying: Pinterest’s $4 Billion Bet
The clearest evidence that this isn’t infrastructure for infrastructure’s sake came from AWS’s disclosure that Pinterest signed a $4 billion AI infrastructure deal with the company, described as Pinterest’s largest infrastructure commitment in its corporate history. Deals at that scale from consumer-facing platforms suggest recommendation systems, ad targeting, and content moderation pipelines are now consuming enough GPU-hours to justify multi-year, multi-billion-dollar commitments rather than on-demand spend.
AWS is also putting new money behind adoption, not just supply. The company is investing $1 billion to embed AI “forward deployed engineers” directly with customer teams, a model that echoes Palantir’s approach, to help enterprises get AI workloads into production rather than stuck in pilot purgatory. It has also launched AWS Secret Cloud for Industry, a faster, more secure path for defense contractors and regulated organizations to build classified AI solutions, and committed more than $500 million to a new Student Rewards program on AWS Builder Center, giving university students worldwide free credits and certification vouchers.
Market Impact: What This Means for Cloud Pricing and Capacity
GPU scarcity has been the single biggest driver of cloud AI pricing since 2023. A 2-million-unit commitment, spread across 2027 and 2028, won’t ease pressure this year, but it signals AWS expects demand to keep climbing well past current levels rather than plateau. For finance and platform teams doing FinOps planning, that has two practical implications. First, expect AWS to keep pushing customers toward reserved and committed-use pricing for GPU instances, since the company is locking in supply years ahead and wants matching demand commitments. Second, expect more regional GPU capacity outside the traditional us-east-1 and us-west-2 hubs, since a buildout this large has to spread across new and expanded data center campuses to avoid concentrating power and cooling risk in a handful of regions.
Amazon’s stock has drawn mixed analyst reaction to the spending pace. One valuation note published the same week flagged that AMZN looked roughly 5.6% overvalued on GF Value metrics amid the accelerating AI infrastructure spend, a reminder that Wall Street is watching capital expenditure against actual AI revenue conversion closely, even as the headline GPU numbers keep growing (source).
AWS vs Azure vs Google Cloud: The Infrastructure Arms Race
AWS isn’t alone in racing to lock down GPU supply. Microsoft has its own multi-year NVIDIA and AMD chip commitments feeding Azure AI Foundry, and Google has been leaning on its in-house TPU line to cut dependence on third-party GPU allocation entirely. The difference in August 2026 is scale and public specificity: AWS put a hard number, 2 million GPUs, and a hard timeframe, 2027-2028, on the record. That level of disclosure is partly a signal to enterprise customers currently weighing multi-cloud AI strategies, and partly a signal to NVIDIA that AWS intends to remain its largest single customer even as Google’s TPU strategy shrinks NVIDIA’s addressable market at the margins.
| Cloud Provider | Primary AI Chip Strategy | Latest Custom Silicon | Disclosed 2026-2028 GPU Commitment |
|---|---|---|---|
| AWS | NVIDIA GPUs + in-house Trainium/Graviton | Graviton5 (GA, Aug 2026) | ~3M NVIDIA GPUs (1M prior + 2M new, 2027-2028) |
| Microsoft Azure | NVIDIA + AMD GPUs + in-house Maia | Maia 200 (in development) | Multi-year, figures not fully disclosed publicly |
| Google Cloud | In-house TPUs + NVIDIA GPUs | TPU v7 (Ironwood, GA 2026) | TPU-first strategy; GPU commitments not publicly quantified |
| Oracle Cloud Infrastructure | NVIDIA GPUs, large Stargate-linked buildouts | No in-house AI chip | Large multi-billion-dollar NVIDIA orders, unit counts undisclosed |
Historical Context: From EC2’s 20th Birthday to the GPU Supercycle
There’s a symbolic backdrop to this announcement: Amazon EC2 turned 20 years old this August, and AWS marked the anniversary on its own blog just days before the NVIDIA news broke. EC2 launched in 2006 as a way to rent raw virtual machines by the hour, a radical idea at the time. Two decades later, the same underlying business, renting compute capacity, has scaled into a market where a single infrastructure commitment involves millions of specialized AI chips and multi-billion-dollar customer contracts.
The shift from general-purpose virtual machines to GPU-dense AI factories mirrors what happened with mobile a decade earlier: the underlying economics moved from renting generic compute cheaply to renting specialized compute at massive scale for a narrower, higher-value workload. AWS’s original EC2 pitch was elasticity for web applications. Its 2026 pitch is elasticity for frontier model training and agentic inference, backed by chip commitments that dwarf the entire compute footprint of early AWS.
Competitive Comparison: GPU Supply as the New Cloud Moat
For most of cloud computing’s history, the competitive moat was breadth of managed services and global region count. In 2026, the moat is increasingly GPU supply and interconnect bandwidth. AWS’s approach layers three things: NVIDIA GPUs for maximum compatibility with existing AI tooling, Trainium and Graviton for cost-optimized workloads AWS controls end to end, and now DuckDB for fast analytical queries feeding those workloads. Google’s approach leans harder into vertical integration through TPUs, which removes NVIDIA dependency but narrows tooling compatibility for customers used to CUDA. Microsoft sits in between, buying heavily from both NVIDIA and AMD while developing Maia chips more slowly.
The practical takeaway for engineering teams choosing a cloud for AI workloads: AWS is betting that flexibility, supporting every major chip and model family, beats vertical integration around one company’s chip and one company’s models. Whether that bet pays off depends on whether GPU supply constraints ease enough by 2027 that flexibility stops being a scarcity hedge and starts being just a feature checkbox.
| Announcement (Aug 2026) | Date | Scale / Detail | Primary Beneficiary |
|---|---|---|---|
| AWS-NVIDIA GPU expansion | Aug 27, 2026 | 2M additional GPUs, 2027-2028 | Enterprise AI training/inference customers |
| AWS acquires DuckLabs (DuckDB) | Aug 25-26, 2026 | Definitive agreement, terms undisclosed | Data engineering and AI agent teams |
| Graviton5 general availability | Mid-August 2026 | 5th-gen Arm chip, agentic-AI positioning | Cost-sensitive general compute and inference |
| Claude Opus 5 / Fable 5 on Bedrock | August 2026 | Frontier model access via Bedrock | AI application developers |
| OpenAI Daybreak Red/Blue on Bedrock | August 2026 | Security-specialized models | Security and vulnerability research teams |
| Pinterest $4B AI infrastructure deal | August 2026 | Largest infra commitment in Pinterest’s history | AWS enterprise sales |
What Developers and Platform Teams Should Watch
None of this changes what’s available to provision today, but it should change planning conversations happening right now for 2027 budgets. Teams running large training jobs on AWS should expect more regional GPU capacity to come online through 2027, which historically correlates with AWS adjusting reserved-instance and Capacity Blocks pricing for P-series and Trn-series instances. Teams building data pipelines should watch for DuckDB’s execution engine showing up as a first-party option inside SageMaker, Glue, or Bedrock Knowledge Bases over the next two to three quarters, since AWS rarely acquires a team without productizing the technology within a year.
Security teams have a more immediate item to evaluate. OpenAI’s Daybreak Red and Daybreak Blue models are now accessible through Bedrock for eligible customers, specifically built for vulnerability discovery and exploit reproduction. That access is governed, not open to every account, but it puts frontier offensive-security tooling inside the same cloud console teams already use for production workloads, which raises new questions about access control and audit logging for any organization that requests it.
H3 Deep Dive: Regional Capacity and Power Constraints
Deploying 2 million GPUs is a power and cooling problem as much as a chip-supply problem. Each high-end NVIDIA accelerator in a dense rack configuration draws well beyond what most legacy data center campuses were built to support, which is why hyperscalers including AWS have been signing long-term power purchase agreements alongside chip orders. Expect the practical rollout of this commitment to track new data center construction and grid interconnection approvals as closely as it tracks NVIDIA’s own production capacity, meaning some regions will see GPU availability improve faster than others depending on local power infrastructure.
H3 Deep Dive: What “Agentic AI” Means for Infrastructure Sizing
AWS repeatedly framed both Graviton5 and the NVIDIA expansion around “agentic AI,” a shift from infrastructure sized for single large training runs to infrastructure sized for millions of smaller, concurrent inference calls made by autonomous agents completing multi-step tasks. That workload pattern favors different infrastructure than frontier training: more inference capacity spread across more regions, lower per-call latency requirements, and tighter integration between compute and fast data access, which is exactly the gap DuckDB is meant to fill.
Predictions: Where This Goes Through 2027
- AWS will formally productize DuckDB’s engine inside at least one existing analytics service, most likely SageMaker or Glue, within two to three quarters, rather than running it as a standalone offering.
- Expect Microsoft and Google to respond with their own large, publicly quantified GPU or TPU capacity announcements before the end of 2026, since AWS setting a hard 2-million-unit number creates competitive pressure to match the disclosure, not just the capacity.
- Reserved and committed-use GPU pricing on AWS will get more aggressive discounting tiers through 2027 as AWS tries to convert this new supply into locked-in multi-year enterprise contracts similar to the Pinterest deal.
- AWS Secret Cloud for Industry will expand beyond defense contractors into other regulated sectors, such as healthcare and financial services, by mid-2027, following the same classified-capacity-as-a-service model.
- Graviton5 adoption will accelerate fastest among inference-heavy workloads rather than training, since AWS is positioning it for cost-efficient agentic AI serving rather than frontier model training, which will keep flowing to NVIDIA GPUs and Trainium.
Frequently Asked Questions
How many NVIDIA GPUs is AWS actually committing to?
AWS announced 2 million additional NVIDIA GPUs for deployment in 2027-2028, on top of an earlier 2026 commitment of more than 1 million GPUs including Blackwell and Rubin architectures. That puts AWS’s total disclosed NVIDIA GPU pipeline at roughly 3 million units.
What is DuckDB and why did AWS buy it?
DuckDB is an open-source, in-process analytical database built for fast queries on datasets under roughly one terabyte, popular with data engineers, data scientists, and increasingly AI agents that need quick structured-data lookups. AWS acquired DuckLabs, the company behind it, to bring that engine in-house rather than depend on a third party.
Is Graviton5 available now?
Yes. AWS confirmed general availability of Graviton5 in mid-August 2026, positioning it as the chip built for the agentic AI era with improved price-performance over prior Graviton generations.
What AI models are new on Amazon Bedrock this month?
Anthropic’s Claude Opus 5 and Claude Fable 5 are both available on Bedrock, alongside OpenAI’s Daybreak Red and Daybreak Blue, specialized models for vulnerability discovery and exploit reproduction available to eligible customers.
Does this affect AWS pricing for GPU instances right now?
Not immediately. The 2-million-GPU deployment runs through 2027-2028, so near-term pricing and availability for P-series and Trn-series instances won’t shift overnight. It does signal AWS expects sustained demand growth and will likely expand reserved and Capacity Blocks pricing options as new regional capacity comes online.
How does this compare to what Microsoft and Google are doing?
Microsoft continues buying heavily from both NVIDIA and AMD for Azure while developing its own Maia chips more slowly. Google leans on its in-house TPU line, now on its seventh generation (Ironwood), to cut NVIDIA dependence. AWS is the only one of the three to put a specific, multi-million-unit NVIDIA GPU number on the record for 2027-2028.
What is AWS Secret Cloud for Industry?
It’s a new AWS offering giving defense contractors and similarly regulated organizations a faster, more secure path to build classified AI solutions, part of the broader government-focused AI factory buildout announced alongside the NVIDIA GPU expansion.
Why did Amazon mention EC2 turning 20 in the same news cycle?
EC2 launched in 2006 as AWS’s original pay-by-the-hour virtual machine service. AWS marked its 20th anniversary in August 2026, days before the NVIDIA expansion, framing the GPU buildout as the next chapter of the same elastic-compute business model that started with EC2.
Related Coverage
- AWS Bedrock AgentCore Hits GA, Cuts AI Costs 80% [2026]
- AWS Bedrock Web Search Goes GA, Google Ships 200+ Models [2026]
- AWS vs GCP vs Azure: GCP Cuts SQL Costs 30% [2026]
- Google Cloud Hits 63% Growth, Outpaces AWS, Azure [2026]
- AWS Outage Hits 28 Hours, Third us-east-1 Failure [2026]
- Kubernetes 1.37 Ships Aug 26: 3 Breaking Changes to Fix Now [2026]




