Google Cloud’s seventh-generation Tensor Processing Unit, code-named Ironwood, has moved from stage demos to a line item on customer invoices, and the pricing tells a story Nvidia would rather not have published. According to Google Cloud’s current pricing documentation, Ironwood instances in the us-central1 region list at $12.00 per chip-hour on demand, with discounted flexible-start options as low as $6.00 per chip-hour. Those numbers alone don’t mean much. What matters is what they translate into on a cost-per-token basis, and recent inference benchmarking puts Ironwood ahead of Nvidia’s B200 and B300 accelerators on exactly that metric. For a company that has spent three years watching Nvidia set the price and the pace of the AI hardware market, that’s a meaningful shift worth unpacking.

This is a story about economics, not about a single chip winning a spec sheet fight. Training FLOPs still matter for frontier model builders, but the bulk of AI infrastructure spending in 2026 is shifting toward serving inference traffic at scale, where cost per million tokens is the number finance teams actually track. Google’s own hardware is now putting pressure on that number in ways that ripple through cloud contracts, GPU rental markets, and the broader hardware cluster competitors have built around Nvidia’s supply.

What Google’s Ironwood TPU Actually Is

Ironwood is Google’s seventh-generation Tensor Processing Unit, built specifically for large-scale training, reasoning workloads, and high-volume inference. Google Cloud describes it as the company’s most capable custom AI accelerator to date, and it has been generally available on Google Cloud since spring 2026. Unlike Nvidia GPUs, which Google and every other hyperscaler can buy, rent out, or resell, Ironwood is Google’s own silicon, designed in-house and deployed exclusively inside Google’s data centers and through Google Cloud’s TPU service. That distinction matters for the pricing conversation: Google controls the entire stack, from chip design to the rate it charges per hour, which is part of why it can undercut Nvidia-based instances on certain workloads.

A full Ironwood pod scales to 9,216 liquid-cooled chips, delivering an aggregate peak of 42.5 exaFLOPS, according to Google Cloud’s published TPU specifications. Google says that works out to roughly four times the per-chip performance of Trillium, the TPU v6e generation that preceded it. That’s a steep generational jump, and it’s the kind of leap Nvidia itself has made with its Blackwell-to-Blackwell Ultra transition, but Google is making the case entirely on its own infrastructure rather than through partner OEMs.

Ironwood’s Memory and Bandwidth Numbers

Memory is where Ironwood’s design choices become clearer. Each chip carries 192GB of HBM3e memory with roughly 7.37 to 7.38 TB/s of bandwidth, figures that put it in the same tier as Nvidia’s current data-center memory configurations rather than a generation behind. Peak FP8 throughput lands at 4,614 TFLOPS per chip, a number Google has been citing in comparisons against both its own prior generation and against Nvidia’s inference-focused configurations. For reference, Trillium (TPU v6e) tops out at roughly 918 BF16 TFLOPS, 32GB of HBM, and 1.638 TB/s of bandwidth per chip, which puts the generational jump in concrete terms: Ironwood ships with six times the memory and more than four times the bandwidth of the chip it replaces.

That memory capacity matters more for inference than it does for raw training throughput. Serving large language models means keeping huge key-value caches resident in fast memory, and running out of HBM capacity is one of the most common reasons inference costs spike at scale. Google’s HBM3e allocation on Ironwood is a direct answer to that bottleneck, and it lines up with the broader industry memory crunch that’s been squeezing Nvidia’s own Rubin Ultra roadmap as HBM supply tightens across the board.

The Pricing Breakdown: What Ironwood Actually Costs

Google Cloud’s published pricing lists three purchase modes for Ironwood capacity in the us-central1 region. On-demand access runs $12.00 per chip-hour. A dynamic workload scheduler flex-start option comes in at $6.00 per chip-hour, half the on-demand rate, aimed at workloads that can tolerate scheduling flexibility. A calendar-mode DWS tier sits in between at $8.40 per chip-hour. Trillium, the prior TPU generation, is listed separately at $2.70 per chip-hour on-demand in the us-east1 region (South Carolina), though that’s a different region than Ironwood’s us-central1 listing, so the two aren’t a perfectly clean apples-to-apples rate comparison.

Here’s where it gets interesting for buyers: a higher sticker price per chip-hour doesn’t necessarily mean a higher cost per unit of useful work, because Ironwood does roughly four times more per chip than Trillium. The real comparison that matters to a company renting capacity to serve a chatbot or an API isn’t the hourly rate, it’s the cost to generate a fixed volume of tokens at a target latency. That’s the number Google and independent benchmarkers have started publishing, and it’s the number that puts real pressure on Nvidia.

gcloud compute tpus tpu-vm accelerator-types list \
  --zone=us-central1-a \
  --filter="type:v7*"

# Check current on-demand and DWS pricing for a given TPU type
gcloud billing accounts list
gcloud alpha billing prices list \
  --filter="service.displayName:'Compute Engine'" \
  --format="table(sku.description, price)"

Ironwood vs Nvidia B200 and B300 on Cost Per Token

Independent inference-cost benchmarking cited in recent industry comparisons looked at two different serving conditions. At an operating point of 100 tokens per second per user, a common proxy for interactive chat-style serving, the reported cost per million tokens came out to $0.181 for Ironwood, $0.222 for Nvidia’s B200, and $0.276 for the newer B300. That puts Ironwood roughly 18% cheaper than B200 and about 34% cheaper than B300 at that specific operating point.

A second comparison, measured at a 20-second median response-time target rather than a fixed tokens-per-second rate, told a similar story with a narrower gap: $0.098 per million tokens for Ironwood versus $0.106 for B200 and $0.132 for B300. That’s roughly an 8% edge over B200 and a 26% edge over B300 under those serving conditions. Both comparisons point the same direction, but the size of Ironwood’s advantage clearly depends heavily on the workload profile, batch size, and latency target being tested, which is worth remembering before treating either number as universal.

This is also showing up against a backdrop of rising Nvidia rental prices elsewhere in the market. Nvidia B200 cloud pricing has climbed to roughly $8.01 per GPU-hour on the open rental market this year, up 79% as demand outstrips available supply. When the GPU side of the ledger is getting more expensive at the same time Google is publishing TPU cost comparisons that favor Ironwood, cloud architects have a reason to at least run the numbers on their own workloads.

The Numbers: Specs and Cost Comparison Tables

Ironwood vs Trillium vs Nvidia B200/B300 Specs

AcceleratorMemory per chipBandwidth per chipPeak throughputOn-demand price
Google Ironwood (TPU v7)192GB HBM3e~7.37-7.38 TB/s4,614 TFLOPS (FP8)$12.00/chip-hr (us-central1)
Google Trillium (TPU v6e)32GB HBM1.638 TB/s918 TFLOPS (BF16)$2.70/chip-hr (us-east1)
Nvidia B200192GB HBM3e~8 TB/s (per Nvidia specs)Varies by precision/config~$8.01/GPU-hr (open rental market)
Nvidia B300288GB HBM3e~8 TB/s (per Nvidia specs)Varies by precision/configHigher than B200 per published cost comparisons

The spec table underscores a point that gets lost in headline comparisons: Ironwood’s memory capacity is now on par with, not behind, Nvidia’s B200. That wasn’t true of prior TPU generations, which typically trailed Nvidia on raw memory and bandwidth even when they competed on price. Closing that gap is arguably more significant than the headline cost-per-token numbers, because it means Google is no longer trading memory capacity for a lower price tag.

Inference Cost Per Million Tokens

Operating pointIronwoodNvidia B200Nvidia B300Ironwood’s edge vs B200
100 tokens/sec/user$0.181 per million tokens$0.222 per million tokens$0.276 per million tokens~18% cheaper
20-second median response time$0.098 per million tokens$0.106 per million tokens$0.132 per million tokens~8% cheaper

Why Cost Per Token Became the Benchmark That Matters

For most of the generative AI boom, the headline hardware metric was raw FLOPs, the kind of number Nvidia used to sell successive generations of H100, H200, and B200 silicon into every major data center buildout. That made sense when the industry was mostly training new frontier models. Training is a one-time capital cost measured in GPU-hours burned over weeks or months. Inference is different. It’s an ongoing operating cost that scales directly with user traffic, and it compounds every single day a product stays live.

As AI products from chatbots to coding assistants to agentic workflows have moved from pilot projects to production traffic serving hundreds of millions of users, the economics of serving inference at scale have overtaken training FLOPs as the number that actually drives procurement decisions. A chip that trains slightly slower but serves tokens meaningfully cheaper can still win a procurement contract, because the serving bill dwarfs the training bill over a product’s lifetime. That’s the argument Google is making with Ironwood, and it’s a direct challenge to the assumption that Nvidia’s CUDA ecosystem and raw compute lead automatically translate into the lowest total cost of ownership.

From TPU v1 to Ironwood: A Decade of Catching Up

Google has been building its own AI accelerators since TPU v1 launched internally in 2015, originally to speed up search ranking and translation workloads rather than to compete with Nvidia on the open market. For years, TPUs were a Google-only tool, available to outside developers only through Google Cloud and mostly used by customers already deep in the Google ecosystem. The TPU v4 and v5 generations narrowed the gap on raw performance but still generally trailed Nvidia’s contemporaneous data-center GPUs on memory capacity and software ecosystem maturity. Trillium (TPU v6e) was the generation that started closing that gap meaningfully, and Ironwood is the first generation where Google is willing to publish head-to-head cost comparisons against Nvidia’s current flagship data-center silicon rather than quietly undercutting it on price alone.

That history matters for context: this isn’t Google suddenly discovering AI chip design. It’s the payoff of roughly a decade of iterative hardware investment finally reaching a point where the silicon is good enough, and the memory capacity large enough, to compete on Nvidia’s own terms rather than just on being the cheaper alternative for customers already locked into Google Cloud.

Market Impact: What This Means for Nvidia’s Position

Nvidia still controls the overwhelming majority of the AI accelerator market, and nothing about Ironwood’s pricing changes that in the short term. Nvidia’s GPUs remain the default choice for any team that isn’t already deep inside Google Cloud, in large part because CUDA and the surrounding software ecosystem are far more portable across cloud providers than TPUs, which are a Google Cloud exclusive. But exclusivity cuts both ways. It means Ironwood’s cost advantage only matters to customers willing to build or already running on Google’s stack, yet for those customers, it is a genuine lever against a market where Nvidia holds something like 90% share of GPU shipments and has largely been able to set its own pricing.

Google isn’t the only hyperscaler pushing custom silicon as a hedge against Nvidia pricing power. Amazon has been scaling its own Trainium chips, Microsoft has its Maia accelerators, and Meta has been designing its own inference silicon. What makes Ironwood notable is that Google is now the first of that group willing to put specific, workload-level cost-per-token numbers next to Nvidia’s current flagship parts in public comparisons, rather than just claiming an advantage in marketing language. That’s a different kind of competitive pressure than a spec sheet war, because it’s the exact number a cloud architect’s finance team will ask for before signing a multi-year capacity commitment.

The Memory Supply Angle: Why Ironwood’s Timing Matters

Ironwood’s 192GB HBM3e allocation per chip is landing at a moment when high-bandwidth memory is one of the tightest components in the entire AI supply chain. DRAM and HBM pricing has been climbing sharply through 2026, and that scarcity has already forced Nvidia to make trade-offs on its own next-generation Rubin platform, where reports indicate Rubin Ultra configurations are losing a meaningful share of planned memory allocation to supply constraints. Google securing enough HBM3e to ship Ironwood at scale, while simultaneously undercutting Nvidia on cost per token, suggests Google’s memory supply agreements are currently in a stronger position than some of Nvidia’s own next-generation plans.

That’s a meaningful signal for anyone tracking the broader AI hardware supply chain. When the chip with the memory supply to back up its launch is also the chip winning cost-per-token comparisons, it points to memory allocation, not raw compute design, being the real constraint shaping who wins the next round of the inference-economics competition.

The Competitive Landscape Beyond Google and Nvidia

Nvidia’s Counter-Moves

Custom silicon from hyperscalers is only part of the pressure building against Nvidia’s pricing. Nvidia’s own Vera CPU platform, which CoreWeave has deployed with 11,264 cores aimed at agentic AI workloads, shows Nvidia is trying to defend its position by expanding beyond GPUs into the CPU side of AI infrastructure rather than just iterating faster on GPU generations. Meanwhile, benchmark comparisons like the Vera Rubin NVL72’s 3.7x lead over GB300 in the first MLPerf Inference v6.1 run show Nvidia still has a substantial raw-performance advantage on its newest platform generation, even as Google chips away at the cost side of the equation on current-generation hardware.

The Rest of the Field

AMD’s Instinct line and Amazon’s Trainium chips add further pressure from different angles, Trainium through AWS’s own cloud exclusivity model that mirrors Google’s approach, and AMD’s MI-series through an open, Nvidia-alternative hardware play that doesn’t require locking into a single cloud provider. None of these alternatives has displaced Nvidia’s market share in any meaningful way yet, but the number of credible cost-competitive alternatives has clearly grown through 2026, and Ironwood’s published cost-per-token numbers are the most concrete evidence yet that at least one of them can win on price for specific inference workloads.

Why This Isn’t an Apples-to-Apples Comparison

A few caveats matter before treating Ironwood’s cost advantage as a universal rule. First, the Ironwood and Trillium list prices come from different Google Cloud regions, us-central1 and us-east1 respectively, and regional pricing can vary independently of the hardware itself. Second, the cost-per-token comparisons against B200 and B300 are workload-specific: they depend on model architecture, batch size, precision, compiler optimization, networking configuration, and whether the comparison measures a single chip, a full node, or an entire pod. Nvidia’s own hardware can post very different cost-per-token numbers depending on how aggressively a given cloud provider tunes its serving stack, and Google has an obvious incentive to publish the comparisons that make Ironwood look best.

None of that invalidates the numbers. It just means buyers should treat published cost-per-token comparisons the way they’d treat any vendor benchmark: directionally useful, but worth validating against your own specific model, traffic pattern, and latency requirements before signing a capacity commitment based on someone else’s test conditions.

What Enterprises and Developers Should Watch Next

Teams already running workloads on Google Cloud have the most immediate reason to test Ironwood against their current Nvidia-based serving stack, since the switching cost is lower when you’re not also migrating cloud providers. Teams running multi-cloud or Nvidia-only infrastructure face a bigger decision: TPUs mean rewriting serving code around Google’s XLA compiler stack and giving up the portability CUDA provides across AWS, Azure, and on-prem deployments. That portability tax is real, and it’s a big part of why Nvidia’s ecosystem lock-in has survived years of cheaper alternative silicon from every major hyperscaler.

Procurement teams should also watch Google Cloud’s regional pricing as Ironwood capacity expands beyond us-central1, since early-generation accelerator pricing often shifts once supply catches up with demand. And anyone benchmarking their own workloads should run the comparison at their actual production latency target rather than relying on either of the two operating points cited in current industry comparisons, since the gap between Ironwood and Nvidia’s chips clearly narrows or widens depending on exactly what’s being measured.

Predictions: Where TPU vs GPU Economics Go From Here

  • Expect Google to expand Ironwood’s regional footprint beyond us-central1 through the rest of 2026, which should bring Trillium-style regional pricing variance into sharper focus for cost comparisons.
  • Nvidia is likely to respond with more aggressive inference-tuned software optimizations for B200 and B300 rather than a price cut, since protecting margin on current-generation silicon matters more to Nvidia than matching Google chip-hour for chip-hour.
  • Other hyperscalers running custom silicon, particularly Amazon’s Trainium and Microsoft’s Maia line, are likely to publish their own cost-per-token comparisons against Nvidia within the next few quarters, following the precedent Google has now set.
  • HBM supply will remain the key swing factor. Whichever platform secures the most favorable long-term memory supply agreements will have the most room to cut inference pricing further without sacrificing margin.
  • Expect enterprise procurement conversations to increasingly ask vendors for cost-per-token figures at a specified latency target as a standard line item, rather than accepting raw FLOPs or chip-hour pricing as the primary comparison metric.

The Bigger Picture for the Hardware Cluster

Ironwood’s cost-per-token numbers land in the middle of a broader hardware story that’s been building through 2026: memory scarcity reshaping chip roadmaps, cloud GPU rental prices climbing even as alternatives multiply, and every major hyperscaler racing to reduce dependence on a single supplier for the most expensive line item in their AI infrastructure budgets. None of that displaces Nvidia from its dominant position this year. What it does is give buyers, for the first time in a while, a credible, numbers-backed reason to at least run the comparison before defaulting to Nvidia GPUs for every inference workload. That’s a meaningful shift in a market that’s spent three years treating Nvidia’s pricing as the only number that mattered.

Frequently Asked Questions

What is Google’s Ironwood TPU?
Ironwood is Google’s seventh-generation Tensor Processing Unit (TPU v7), a custom AI accelerator built for large-scale training, reasoning, and high-volume inference. It has been generally available on Google Cloud since spring 2026.

How much does Google’s Ironwood TPU cost?
Google Cloud lists Ironwood at $12.00 per chip-hour on demand in the us-central1 region, with a flex-start dynamic workload scheduler rate of $6.00 per chip-hour and a calendar-mode rate of $8.40 per chip-hour.

Is Ironwood cheaper than Nvidia’s B200 and B300?
On cost per million tokens for inference, published comparisons show Ironwood coming in cheaper than both Nvidia chips at two tested operating points, roughly 18% cheaper than B200 and 34% cheaper than B300 at a 100 tokens/sec/user target, and a narrower 8% and 26% advantage respectively at a 20-second median response-time target. Results vary by workload.

How much memory does Ironwood have per chip?
Each Ironwood chip ships with 192GB of HBM3e memory and roughly 7.37 to 7.38 TB/s of memory bandwidth, putting it on par with Nvidia’s B200 memory capacity.

Can I use Ironwood TPUs outside of Google Cloud?
No. Unlike Nvidia GPUs, which are sold to and deployed by many different cloud providers and enterprises, TPUs including Ironwood are exclusive to Google’s own infrastructure and are only accessible through Google Cloud.

How big is a full Ironwood pod?
A full Ironwood pod scales to 9,216 liquid-cooled chips, delivering an aggregate peak of 42.5 exaFLOPS, according to Google Cloud’s published specifications.

Does switching to TPUs require rewriting my AI serving code?
Generally yes. TPU workloads run through Google’s XLA compiler stack rather than Nvidia’s CUDA ecosystem, so migrating an existing CUDA-based serving pipeline to TPUs typically requires engineering work and testing before production deployment.

Will Nvidia cut GPU prices in response to Ironwood?
There’s no confirmation of a price response from Nvidia. Given current GPU rental prices have been rising rather than falling amid strong demand, a near-term Nvidia price cut specifically tied to TPU competition looks unlikely; a software-performance response is more probable.