Renting an Nvidia H100 for an hour costs anywhere from $1.49 to nearly $13 in September 2026, depending on who you ask and which cloud you pick. That spread, once a rounding error in a niche market, has become one of the clearest signals of how tight the AI compute supply chain still is, two years after ChatGPT turned GPU capacity into a boardroom talking point. Pricing trackers, bank analysts, and Nvidia’s own executives are now telling three overlapping stories at once: H100 rental costs have fallen by roughly half since early 2024, newer chips like the H200 and B200 command a steep premium, and the entire market remains bottlenecked by memory, not by silicon.
The numbers matter beyond the spreadsheets of infrastructure buyers. Every dollar per GPU-hour ripples into the cost of running a large language model, training a new frontier system, or serving a single chatbot query. This piece breaks down where H100, H200, and B200 rental prices actually stand today, why they diverge so sharply between marketplaces and hyperscalers, what Nvidia’s leadership has said about the supply gap, and how AMD and Google’s chips stack up as alternatives.
What’s Driving the H100, H200, and B200 Price Swings Right Now
Three forces are colliding in the cloud GPU market this month. First, demand for inference capacity keeps climbing as more companies ship AI features into production, not just experiment with them. Second, the memory supply chain has tightened sharply, with high-bandwidth memory and DRAM both running short as manufacturers redirect capacity toward data-center chips. Third, Nvidia itself has said it cannot fully meet demand even with record shipments, which pushes buyers toward secondary marketplaces and specialist clouds when hyperscaler capacity runs out.
Pricing tracker Thunder Compute, in a blog post on H100 pricing dated September 1, 2026, lists on-demand H100 80GB rates spanning from $2.89 an hour on RunPod up to $10.98 an hour on Google Cloud’s a3-highgpu-8g instances. CloudZero’s own H100 cost breakdown, published in late August 2026, puts the full spread even wider: Vast.ai as low as $1.49 an hour, Azure as high as $6.98 per GPU-hour on demand. Between those two data points sits a real number that infrastructure teams have to budget against every quarter.
Cloud GPU Rental Prices in September 2026: The Numbers
Three tiers have emerged in the market and they behave almost like separate products. Marketplace and community-cloud rates sit at the floor, where platforms like RunPod’s community tier and Vast.ai’s peer-to-peer listings undercut everyone else. Specialist clouds such as Lambda, Nebius, and Crusoe occupy the middle, offering more reliability and support without hyperscaler markup. Hyperscaler on-demand pricing from AWS, Azure, and Google Cloud sits at the top, often two to four times the marketplace floor for the exact same chip.
| Chip | Marketplace / community floor | Specialist cloud (mid-tier) | Hyperscaler on-demand |
|---|---|---|---|
| H100 80GB | $1.49–$2.99/hr | $3.19–$4.59/hr | $6.88–$10.98/hr |
| H200 141GB | $3.59/hr (RunPod Community) | $4.31–$4.59/hr | $5.58–$6.35/hr |
| B200 | Limited public data | $5.89/hr (RunPod tier cited by Hackceleration) | $8.64/hr (normalized Flex estimate) |
Sourcing for that table pulls from Thunder Compute’s H200 pricing tracker, RunPod’s own published rate cards, and CloudZero’s cross-cloud comparison. B200 numbers remain the least reliable of the three chips. Nvidia has shipped Blackwell in volume since early 2026, but most cloud providers still quote B200 capacity through sales conversations rather than public rate cards, which is why independent trackers keep flagging the chip as data-scarce even six months after general availability.
Nvidia H100 Pricing: From Marketplace Floors to Hyperscaler Premiums
The H100 is now three years old as a product line, which explains why it has the deepest, most consistent pricing data of any chip in this market. Bank of America analyst Vivek Arya, in a research note covered by TheStreet on August 12, 2026, put H100 spot rental prices at roughly $2.80 an hour, describing GPU spot prices broadly as sitting near all-time highs even for older silicon. A separate Bank of America report cited by TechFlowPost on August 18, 2026 pegged H100 at $2.77 an hour, up 2% month over month and up 33% year over year, while noting that A100 pricing held flat at $1.65 an hour, up 17% from a year earlier.
That year-over-year climb on an aging chip is the detail worth sitting with. Normally, GPU rental prices fall as a chip ages and newer silicon takes over premium workloads. Instead, H100 pricing rose through most of 2026 even as H200 and Blackwell chips became available, because demand for any working GPU capacity, new or old, has outpaced the industry’s ability to add supply.
Nvidia H200: The 141GB Premium Tier
The H200 carries 141GB of HBM3e memory versus the H100’s 80GB, and cloud providers price that memory bump directly into the rental rate. RunPod’s community-cloud listing for H200 141GB has held at roughly $3.59 an hour through most of 2026, according to multiple pricing aggregators including DeployBase and Flexprice, while RunPod’s secure-cloud tier for the same chip runs $4.39 to $4.59 an hour. Hivenet’s own pricing page, updated August 28, 2026, lists H200 SXM configurations as high as $4.31 an hour for cluster-grade deployments.
What stands out is how narrow the gap has become between H100 and H200 pricing at the top end. A specialist-cloud H100 at $4.59 an hour and an H200 at the same price point mean buyers increasingly default to the newer chip whenever the rate cards converge, since the extra memory headroom costs nothing extra at that tier. That dynamic is quietly pushing H100 demand down toward the marketplace floor, which is part of why sub-$2 H100 listings have become common on peer-to-peer platforms this year.
Nvidia B200 and Blackwell Ultra: Why Public Data Stays Thin
Blackwell-generation B200 chips reached general availability across major clouds in the first half of 2026, but transparent, per-hour rental pricing has been slow to show up on public rate cards. The clearest figures come from Bank of America’s tracking: Vivek Arya’s note put B200 spot pricing around $5.66 an hour as of mid-August 2026, while the TechFlowPost-covered BofA report from a week later listed B200 at $5.63 an hour, down 2% month over month but still up 7% from a year earlier.
Retail-style pricing trackers show a wider range. One RunPod pricing guide cited a $5.89-an-hour rate for B200, while a separate normalized comparison put the effective per-hour cost as high as $8.64 once accounting for multi-GPU pod configurations. Blackwell Ultra and the GB200 rack-scale system have essentially no public per-hour pricing at all right now, since most of that capacity is still sold through direct contracts with hyperscalers and large AI labs rather than listed on open marketplaces.
How Prices Have Moved Since 2024: The Chipflation Story
AIMultiple’s Cloud GPU Rental Price Index, updated August 25, 2026, found that the median H100 rental price across its tracked cohort now sits around $3.38 an hour, down from more than $7 an hour in early 2024. That is a genuine price collapse of over half in just two and a half years, and it reflects the normal pattern for an aging chip: more supply comes online, more providers compete for the same customers, and rates fall.
But that story only applies to the marketplace floor. Reuters reported on June 3, 2026 that Morgan Stanley analysts were warning about AI “chipflation,” pointing to memory chip prices that had spiked roughly six-fold over the prior year as manufacturers redirected DRAM and HBM production toward higher-margin data-center parts. A separate Morgan Stanley note covered by CNBC on January 20, 2026 forecast DDR4 pricing climbing as much as 93% to 98% quarter over quarter in the first quarter of 2026 alone. Put simply, chip rental prices have fallen even as the memory inside those chips gets dramatically more expensive to produce, and that gap is starting to show up in hyperscaler list prices for the newest hardware.
| Period | Median H100 rate | Source |
|---|---|---|
| Early 2024 | $7+/hr | AIMultiple Cloud GPU Rental Price Index |
| Mid-2026 | ~$3.38/hr | AIMultiple Cloud GPU Rental Price Index (Aug 25, 2026) |
| August 2026 (spot, BofA) | $2.77–$2.80/hr | Bank of America analyst Vivek Arya, via TheStreet and TechFlowPost |
Nvidia’s Own Words: What Huang and Kress Are Telling Investors
Nvidia’s leadership has been unusually direct about the supply gap on recent earnings calls. During the company’s fiscal Q3 2026 call on November 19, 2025, CFO Colette Kress told analysts, according to a transcript covered by Business Insider, “The clouds are sold out,” adding that Nvidia’s installed base across Blackwell, Hopper, and Ampere generations was “fully utilized.” That comment set the tone for the following two quarters.
By the company’s fiscal Q2 2027 call on August 26, 2026, CEO Jensen Huang went further, telling investors, per a transcript cited by Fortune and MarketBeat, “Our supply allows us to confidently deliver 70% growth, though demand is much higher,” and separately that “our entire supply chain is challenged… we have supply for 70% growth, but demand is much higher.” Kress followed with a warning that stuck with analysts: supply would “remain a bottleneck, at least through the end of fiscal year 28,” according to Investopedia’s live coverage of the call. TechTimes reported that same day that Kress described customer forecasts pointing to demand roughly doubling the following year, well beyond what Nvidia’s guidance assumes it can physically ship.
The Memory Bottleneck Behind the Price Floor
Every executive statement above points to the same root cause: memory, not GPU die output, is the binding constraint. Morgan Stanley’s note on Micron and SanDisk, summarized by Yahoo Finance on June 3, 2026, described DRAM as having become “the principal bottleneck in the AI infrastructure buildout,” with the bank modeling DRAM pricing up 40% in the May 2026 quarter and a further 15% in August. Hyperscalers, the note said, remain willing to pay those elevated prices rather than slow down deployments.
That memory squeeze explains why B200 and H200, both of which carry more and faster HBM than the H100, show the least public pricing transparency and the steepest premiums in this market. It also explains why older chips like the A100 and H100 have held or gained value instead of depreciating on schedule. When the newest chip’s memory bill goes up, every chip below it in the stack becomes relatively more attractive, which keeps demand for aging hardware higher than it would otherwise be.
AMD Instinct MI300X and MI355X: The Discount Alternative
AMD’s Instinct line has carved out a real position as the lower-cost alternative to Nvidia’s stack, and the pricing data backs that up. Thunder Compute’s MI300X tracker, updated September 1, 2026, lists RunPod at $2.39 an hour, DigitalOcean at $2.59 an hour, Crusoe Cloud at $3.45 an hour, and Azure at $6.00 to roughly $7.86 per GPU-hour on its ND96isr_MI300X_v5 instance. DeployBase’s own comparison from late August 2026 found DigitalOcean pricing as low as $1.99 an hour, with HotAisle offering single-GPU MI300X access at the same rate.
MI355X, AMD’s newer chip, has almost no public pricing history yet. GPUPerHour’s rental tracker, last updated July 30, 2026, found exactly one provider listing the chip, at $2.59 per GPU-hour, which underscores how early MI355X still is in its cloud rollout compared to Nvidia’s more mature H200 and B200 listings.
Google Cloud TPU v6e and v7: A Different Pricing Model Entirely
Google prices its Trillium TPU v6e chips on its own published pricing page, which lists on-demand access in the U.S. East region at $2.70 per chip-hour, with discounted tiers available: $1.35 per chip-hour under a flex-start commitment, $1.89 under a one-year commitment, and $1.22 under a three-year commitment. That published, tiered structure is unusual in this market, where most Nvidia and AMD rental pricing comes from third-party trackers rather than the chipmaker itself.
Ironwood, Google’s newer v7 TPU, has no public on-demand rate at all as of early September 2026. Analyst estimates cited by TECHi and Spheron point to a negotiated capacity rate near $1.60 per TPU-hour for Anthropic specifically, a figure that reflects a large custom deal rather than a retail price. Other trackers, including one from research firm Margins, estimate Ironwood on-demand pricing could run as high as $12 per chip-hour once Google eventually publishes a public rate, underlining just how much guesswork still surrounds pricing for the newest AI chips across every vendor, not just Nvidia.
Competitive Comparison: Nvidia vs AMD vs Google on Cost per Hour
| Vendor / chip | Cheapest tracked rate | Typical mid-tier rate | Notes |
|---|---|---|---|
| Nvidia H100 | $1.49/hr (Vast.ai) | $3.38/hr (median, AIMultiple) | Down from $7+/hr in early 2024 |
| Nvidia H200 | $3.59/hr (RunPod Community) | $4.31–$4.59/hr | 141GB HBM3e, narrowing gap with H100 |
| Nvidia B200 | $5.63–$5.66/hr (BofA spot) | $5.89–$8.64/hr | Thin public data six months post-launch |
| AMD MI300X | $1.85/hr (Vultr) | $2.39–$3.45/hr | Consistently undercuts equivalent Nvidia tiers |
| AMD MI355X | $2.59/hr (single provider) | Not enough data | Very early cloud availability |
| Google TPU v6e (Trillium) | $1.22/hr (3-yr commit) | $2.70/hr on-demand | Only vendor with a fully public rate card |
| Google TPU v7 (Ironwood) | ~$1.60/hr (Anthropic contract est.) | No public on-demand rate | Estimates range up to $12/hr |
The pattern across every vendor is consistent: newer, higher-memory chips carry a premium, and that premium tracks the underlying cost of HBM rather than any change in raw compute performance. AMD remains the most reliable discount option for teams that can tolerate a smaller software ecosystem than Nvidia’s CUDA stack, while Google’s TPU line is the only platform with genuinely transparent published pricing, a structural advantage that Nvidia and AMD resellers have not matched.
Market Impact: Who Wins and Who Pays
Hyperscalers absorb the least pain here. AWS, Azure, and Google Cloud can spread memory cost increases across enormous customer bases and long-term reserved-instance contracts, which is why their on-demand list prices move slower than marketplace spot rates. Specialist clouds like CoreWeave and Nebius sit in the middle, benefiting from the pricing dynamics that Bank of America flagged in an August 17, 2026 note covered by Seeking Alpha, which found both companies “taking advantage of stronger pricing dynamics” as demand for reliable, mid-tier capacity holds up.
Startups and independent developers feel the squeeze hardest. Marketplace floors around $1.49 to $2 an hour for an H100 sound cheap in isolation, but that capacity is often oversubscribed, inconsistent, and unsuitable for production training runs that need guaranteed uptime. Teams that need reliability end up pushed toward the $4-to-$7 specialist tier or the $7-to-$11 hyperscaler tier, a jump of two to four times the headline marketplace rate advertised in most pricing comparisons.
Historical Context: How the Market Got Here
The GPU rental market barely existed in its current form before 2023. ChatGPT’s public launch in late 2022 triggered a scramble for training and inference capacity that caught cloud providers flat-footed, and H100 rental rates north of $7 an hour through 2023 and into early 2024 reflected genuine scarcity rather than deliberate pricing power. As more capacity came online through 2024 and 2025, competition among specialist clouds like Lambda, CoreWeave, and a wave of newer entrants such as RunPod and Vast.ai pulled marketplace prices down toward today’s sub-$2 floor.
What changed in 2026 is that the bottleneck moved one layer down the stack, from GPU die availability to the memory that sits next to it. Nvidia can apparently manufacture enough Blackwell and Hopper silicon to meet a large share of demand, per Huang’s own comments about supply commitments extending into 2027, but HBM and DRAM output has not kept pace, which is why the newest, most memory-hungry chips show the thinnest public pricing and the steepest premiums of the entire product stack.
Five Predictions for GPU Pricing Through 2027
- H100 marketplace rates likely keep drifting toward $1 an hour by mid-2027 as the chip ages further and B200 capacity absorbs premium demand, following the same depreciation curve the A100 followed before it.
- B200 and Blackwell Ultra pricing should become far more transparent over the next two to three quarters as more cloud providers publish standard rate cards, following the pattern H100 and H200 already went through.
- Memory costs, not GPU die supply, will keep setting the pricing floor for new chips through at least fiscal 2028, based on Colette Kress’s own guidance to investors.
- AMD’s MI355X pricing data should firm up quickly once more providers list the chip, likely landing between MI300X’s current rate and Nvidia’s H200 tier.
- Google is likely to publish an official Ironwood on-demand rate within the next two to three quarters, closing the gap between its transparent TPU v6e pricing and the current guesswork around v7.
Frequently Asked Questions
How much does it cost to rent an Nvidia H100 in September 2026?
Rates range from about $1.49 an hour on marketplaces like Vast.ai up to $10.98 an hour on hyperscaler on-demand instances, with a tracked median around $3.38 an hour according to AIMultiple’s cloud GPU rental index.
Is the H200 always more expensive than the H100?
Usually, but the gap has narrowed. RunPod’s community-cloud H200 rate of roughly $3.59 an hour sits close to mid-tier H100 pricing, which is pushing more buyers toward H200 whenever prices converge.
Why is B200 pricing so hard to find?
Most B200 capacity is still sold through direct sales conversations with hyperscalers rather than listed on public rate cards, six months after general availability. Bank of America’s analyst tracking currently offers the most consistent public spot-price figures, around $5.63 to $5.66 an hour.
Is AMD’s Instinct MI300X actually cheaper than Nvidia’s H100?
Yes, on most trackers. MI300X rates as low as $1.85 to $1.99 an hour on Vultr and DigitalOcean undercut equivalent H100 listings, though Nvidia’s CUDA software ecosystem remains more mature than AMD’s ROCm stack for many workloads.
Why has H100 rental pricing risen even though the chip is three years old?
Demand for any working GPU capacity, old or new, has outpaced supply. Bank of America tracked H100 spot pricing up 33% year over year in an August 2026 note, even as newer chips became available.
What is causing the current memory shortage behind these prices?
Manufacturers have redirected HBM and DRAM production toward higher-margin data-center chips. Morgan Stanley described memory prices as having risen roughly six-fold over the prior year in a note covered by Reuters on June 3, 2026.
Does Google publish official TPU pricing?
Yes, for TPU v6e (Trillium), Google lists on-demand and committed-use rates directly on its cloud pricing page. TPU v7 (Ironwood) has no public on-demand rate yet, only analyst estimates.
When will GPU rental prices stabilize?
Nvidia’s CFO Colette Kress told investors on the company’s August 26, 2026 earnings call that supply would likely remain a bottleneck at least through the end of fiscal 2028, suggesting no near-term return to the looser pricing seen in 2024 and 2025.




