Renting a Nvidia B200 by the hour now costs more per unit of AI compute than renting an older H100, and the gap is widening instead of closing. According to a September 25, 2026 market analysis published on the SemiAnalysis GPU pricing index, B200 cloud-rental rates climbed 79% over three months to reach $8.01 per GPU-hour. That single figure captures a pricing pattern that breaks with how GPU markets normally behave: newer chips are supposed to get cheaper per unit of useful work as supply ramps, not more expensive.
The shift matters well beyond the spreadsheets of cloud-compute brokers. It touches every AI startup renting GPUs by the hour, every hyperscaler budgeting 2027 capital expenditure, and every developer wondering why access to frontier-class compute keeps getting harder to plan around. This is a look at what the numbers show, why memory capacity has become the real price-setter, and what happens next as Blackwell-generation chips like the B200 collide with a global high-bandwidth memory shortage that has already reshaped pricing across Nvidia’s own Rubin Ultra roadmap.
Nvidia B200 Cloud Rental Prices Jump 79% in Three Months
The core data point comes from the SemiAnalysis GPU pricing index, which tracks day-to-day cloud rental rates across a basket of providers. As of late September 2026, that index put B200 rental pricing at $8.01 per GPU-hour, a 79% climb from where it sat three months earlier. The same analysis calculated that a B200 now costs 21% more per unit of compute than an H100, a full reversal from June 2026, when a B200 was actually 18% cheaper than an H100 on the same compute-adjusted basis.
That reversal is the headline. Six months ago, buying into the newest Blackwell-generation hardware looked like the economically rational move: more performance per dollar than sticking with Hopper-era H100 chips. Today, the math has flipped, and it flipped fast. A rental decision made in June would have favored the B200. The identical decision made in September favors the H100, at least on a pure compute-cost basis, even though the H100 is now two full generations behind Nvidia’s current lineup.
H200 pricing sits in between, and its position tells its own story. The SemiAnalysis figures put H200 rental rates at roughly 1.76 times H100 pricing, a ratio that tracks closely with the two chips’ relative memory capacity rather than their raw compute throughput. That detail is the thread running through this entire pricing episode.
Nvidia B200 vs H200 vs H100: The Pricing Breakdown
Putting the three chips side by side makes the memory-driven pricing pattern easier to see. The table below combines the cloud-rental figures from the SemiAnalysis index with indicative hardware pricing reported in a separate September 2026 hardware cost analysis.
| GPU | Generation | HBM Memory | Cloud Rental (per GPU-hour) | Indicative Hardware Price | Approx. Cost per GB |
|---|---|---|---|---|---|
| H100 | Hopper | 80 GB HBM3 | Baseline (index reference) | ~$31,000 | ~$387.50/GB |
| H200 | Hopper | 141 GB HBM3e | ~1.76x H100 rate | ~$39,999 | ~$283.68/GB |
| B200 | Blackwell | 192 GB HBM3e | $8.01/hr (+21% vs H100 compute-adjusted) | ~$50,000–$70,000 | ~$260–$365/GB |
Two things stand out. First, per-gigabyte pricing across all three chips has converged into a fairly tight band, somewhere around $260 to $390 per gigabyte of HBM, regardless of which generation the memory sits on. Second, the B200’s wide hardware price range ($50,000 to $70,000) reflects a supply-constrained market where list price means less than what a buyer can actually secure, and when.
Why Memory Capacity Is Now Setting the Price, Not Compute
For most of the last decade, GPU pricing tracked compute throughput. Buyers paid for FLOPS, and memory was a secondary spec that mattered mostly for gaming cards and a narrow set of scientific workloads. That model has broken down for AI accelerators, and the September 2026 pricing data is the clearest evidence yet.
Large language models with long context windows need memory to hold the key-value cache during inference. Bigger batch sizes, needed to keep GPU utilization high and serve more users per chip, also eat memory before they eat compute. A model that technically fits on an H100’s 80 GB of HBM3 might only run a handful of concurrent requests before hitting a memory wall, while the same model on a B200’s 192 GB has room to serve far more traffic per chip. That gap is why renters are willing to pay a premium for memory capacity even when the underlying compute-per-dollar math looks worse.
The SemiAnalysis analysis frames this directly: the normal GPU depreciation curve, where each new generation delivers more performance per dollar than the last, has broken down because newer accelerators aren’t consistently getting cheaper per unit of usable compute. Memory-bound workloads change the unit of measurement entirely, and once memory becomes the constraint, the chip with more of it commands a premium regardless of how the raw compute math shakes out.
The HBM Supply Chain Behind the Numbers
None of this happens in a vacuum. High-bandwidth memory production is concentrated among a small number of suppliers, and 2026 has been a year of persistent HBM tightness across the entire AI hardware stack. Nvidia’s own next-generation Rubin Ultra platform has reportedly had to trim memory allocation per package because of the same shortage, and cloud providers including Nebius have already raised prices to reflect higher input costs on memory-heavy configurations. When the memory that feeds B200 and H200 packages is scarce and expensive to source, that cost gets passed straight through to the hourly rental rate.
Hardware List Price vs Cloud Rental: A Widening Gap
Buying outright and renting by the hour are two very different bets right now, and the numbers show why. A buyer who can secure a B200 at the low end of its reported $50,000–$70,000 range and run it continuously will eventually amortize that cost below the current $8.01/hour cloud rate. But “eventually” is doing a lot of work in that sentence, since availability, not price, is the real bottleneck for direct purchases.
| Metric | H100 (Hopper) | H200 (Hopper) | B200 (Blackwell) |
|---|---|---|---|
| Memory | 80 GB HBM3 | 141 GB HBM3e | 192 GB HBM3e |
| Relative availability (Sept. 2026) | Widest | Moderate | Most constrained |
| Price trend vs. 3 months ago | Reference point | Tracking memory ratio | +79% cloud rental |
| Compute-cost vs. H100 (Sept. 2026) | 1.0x (baseline) | ~1.76x memory-weighted | +21% more expensive |
| Compute-cost vs. H100 (June 2026) | 1.0x (baseline) | n/a | 18% cheaper |
The direction of travel across three months is the story: B200 pricing moved from a discount to a premium relative to H100, a full swing of roughly 39 percentage points in compute-adjusted terms. That is not a rounding error. It is a market repricing a scarce resource in real time.
Historical Context: How GPU Pricing Used to Work
It’s worth remembering what “normal” looked like before the AI boom rewired the GPU market. Through most of the 2010s and into the early 2020s, each new Nvidia data-center generation, from Volta to Ampere to Hopper, arrived at a price premium over its predecessor but delivered enough of a performance jump that cost-per-FLOP kept falling. Buyers who waited a generation got a better deal in raw economic terms almost every time.
The Hopper-to-Blackwell transition broke that pattern for the first time at scale. Nvidia’s own DGX B200 platform page touts the generational leap in FP4 inference throughput and memory bandwidth, and on paper Blackwell chips like the B200 do outperform Hopper’s H200 on raw specs. But paper specs assume a buyer can get the chips in volume at a stable price, and 2026’s supply constraints have made that assumption fall apart. This is also the year Nvidia’s overall AI hardware dominance came under more direct competitive pressure than at any point since Hopper’s launch, with rivals racing to field alternatives and Nvidia’s own desktop GPU shipment dominance drawing fresh scrutiny even as its data-center chips remain the default choice for frontier AI labs.
Market Impact: Who Actually Pays the Higher Bill
The immediate losers in a pricing environment like this are the mid-size AI labs and startups that rent compute by the hour rather than buying hardware outright or signing long-term capacity contracts. A startup fine-tuning a model or serving inference at scale on rented B200s is now paying a premium that didn’t exist three months ago, with no corresponding jump in the compute it’s getting for that money.
Cloud providers themselves are caught in the middle. Neoclouds like CoreWeave and Lambda Labs, which built businesses around reselling Nvidia capacity, have to pass through higher acquisition costs or eat margin. Some have already chosen the former: Nebius, one of the higher-profile AI-cloud operators, raised its own pricing in 2026 citing rising memory costs, a move that lines up directly with the HBM scarcity driving the B200 and H200 numbers here.
Hyperscalers with in-house chip programs are the one group somewhat insulated from this specific dynamic. Google’s TPUs and Amazon’s Trainium chips don’t show up in the SemiAnalysis Nvidia-focused index at all, and that’s precisely the point: owning your own silicon supply chain means less exposure to Nvidia’s memory-driven repricing, even if the same underlying HBM shortage still touches that supply chain from a different angle.
Competitive Landscape: Neoclouds vs Hyperscalers vs In-House Silicon
Three distinct buyer strategies are playing out simultaneously in response to this pricing environment, and each carries different risk.
Neoclouds Betting on Nvidia Supply
Companies whose entire business model is reselling Nvidia GPU capacity by the hour are the most exposed to swings like the one described here. Their margins compress every time acquisition costs rise faster than what customers are willing to pay per GPU-hour, and the September 2026 numbers suggest that squeeze is real and current, not theoretical.
Hyperscalers Hedging With Custom Silicon
Google, Amazon, and Microsoft have all invested heavily in custom AI accelerators precisely to avoid full dependence on Nvidia’s pricing and supply decisions. That hedge looks increasingly smart given how quickly compute-adjusted B200 pricing swung against renters this year.
AMD, meanwhile, has been pushing its own MI-series accelerators as a lower-cost alternative to Nvidia’s Hopper and Blackwell lines, while simultaneously making moves to strengthen its position in the broader AI stack, including its announced acquisition of World Labs. Whether AMD’s silicon can meaningfully dent Nvidia’s pricing power in the near term remains an open question, but every dollar of premium on a B200 rental is a dollar of incentive for buyers to at least evaluate the alternative.
The Ripple Effect on Consumer and Enterprise Hardware
Data-center GPU scarcity doesn’t stay contained to data centers. The same HBM supply chain that feeds B200 and H200 production also feeds Nvidia’s consumer lineup, and 2026 has seen consumer-grade cards swept up in the same dynamic. Nvidia’s flagship gaming card has effectively disappeared from US retail shelves this year, with resale prices spiking well beyond launch MSRP, a pattern that mirrors the data-center story: when the underlying memory supply is tight, every product built on it gets more expensive and harder to find, regardless of what segment it’s aimed at.
The knock-on effects extend into geography as well. Reports throughout 2026 have documented AI chip price increases inside China tied to the same global HBM shortage, underscoring that this is not a US-specific or Nvidia-specific phenomenon. It’s a global memory bottleneck expressing itself through every product category that depends on advanced HBM packaging.
What This Means for AI Startups and Developers Right Now
For teams actually shipping AI products, the practical takeaway is that hourly cloud-GPU budgets built three or six months ago are probably out of date. A team that priced out a B200-based inference deployment in June, using the 18%-cheaper-than-H100 figure, would be underestimating costs today by a swing of nearly 40 percentage points on a compute-adjusted basis.
That has pushed some engineering teams toward a more memory-conscious approach to model deployment: quantization to cut memory footprint, more aggressive batching to spread fixed memory overhead across more requests, and in some cases a deliberate move back to H100 capacity for workloads that don’t strictly need B200-class memory headroom. None of these are new techniques, but the economic incentive to use them has sharpened considerably since June.
Predictions: Where B200, H200, and H100 Pricing Goes Next
Based on the trajectory in the data and the broader HBM supply picture, a few outcomes look more likely than others heading into 2027.
First, the B200 premium over H100 likely persists into early 2027 rather than correcting quickly, since HBM3e supply additions take quarters to come online and won’t catch up with demand overnight.
Second, expect more cloud providers to follow Nebius’s lead and adjust list pricing upward rather than absorb rising acquisition costs, particularly smaller neoclouds without the balance sheet to eat margin compression.
Third, per-gigabyte memory pricing, not per-chip or per-FLOP pricing, becomes the metric more buyers actually negotiate around, since it has proven to be the more stable and predictive figure across all three chip generations in the September 2026 data.
Fourth, competitive pressure from AMD’s MI-series and hyperscaler custom silicon likely intensifies as buyers look for any lever to escape Nvidia-specific pricing swings, even if switching costs keep most workloads on Nvidia hardware for now.
Fifth, watch for Nvidia’s next-generation Rubin platform to inherit this same memory-driven pricing pattern rather than escape it, given that HBM scarcity is already showing up in reduced memory allocations on Rubin Ultra ahead of its own launch.
How to Read GPU Pricing Indexes Going Forward
One practical lesson from this episode is that headline hardware prices tell you less than compute-adjusted and memory-adjusted pricing does. A B200 costing more upfront than an H100 has always been expected. A B200 costing more per unit of usable compute than an H100, for the first time in a Nvidia generational transition, is the actual news. Anyone budgeting AI infrastructure spend into 2027 should be tracking indexes like the ones from SemiAnalysis and ORNN’s compute price index rather than relying on list prices alone, since list prices increasingly don’t reflect what buyers actually pay in a supply-constrained market.
Independent trackers including GPUSmith’s GPU price index and Spheron’s cloud pricing comparison have documented similar patterns across a wider set of providers, reinforcing that the SemiAnalysis figures aren’t an outlier specific to one data source but a broader market signal.
Frequently Asked Questions
Why is the Nvidia B200 more expensive to rent than the H100 right now?
Because memory capacity, not raw compute, has become the primary price driver for AI accelerators in 2026. The B200’s larger 192 GB HBM3e pool commands a premium in a market where high-bandwidth memory is scarce, even though it’s a newer and more powerful chip overall.
How much did B200 cloud rental prices rise in 2026?
According to the SemiAnalysis GPU pricing index, B200 rental rates rose 79% over three months, reaching $8.01 per GPU-hour by late September 2026.
Is it cheaper to buy or rent a B200 in 2026?
It depends on utilization and availability. Indicative purchase prices for a B200 run $50,000–$70,000, but B200 hardware is harder to source in volume than H100 or H200, which affects the real-world cost of the “buy” option regardless of list price.
What is causing the Nvidia HBM memory shortage?
Global demand for AI accelerators has outpaced high-bandwidth memory production capacity, a shortage that has also affected Nvidia’s upcoming Rubin Ultra platform and pushed cloud providers like Nebius to raise their own prices.
How does H200 pricing compare to H100 and B200?
H200 rental prices run about 1.76 times H100 rates, a ratio that closely tracks the two chips’ relative memory capacity (141 GB versus 80 GB) rather than their compute difference.
Will B200 prices come down in 2027?
Not immediately, based on current supply trends. HBM3e production increases take multiple quarters to reach the market, so the current premium is likely to persist into at least early 2027.
Are AMD GPUs a viable alternative to Nvidia’s B200 or H200?
AMD’s MI-series accelerators are positioned as lower-cost alternatives, and rising Nvidia rental costs increase the incentive for buyers to evaluate them, though switching costs and software ecosystem lock-in keep most large-scale AI workloads on Nvidia hardware for now.
What should AI startups do about rising GPU rental costs?
Reassess infrastructure budgets set before September 2026, prioritize memory-efficient techniques like quantization and larger batching, and track compute-adjusted and per-gigabyte pricing rather than sticker price when choosing between H100, H200, and B200 capacity.




