Nvidia is testing stripped-down memory configurations for its next AI accelerator, Rubin Ultra, as a global shortage of high-bandwidth memory forces a rethink of plans headed into 2027. A report published August 4, 2026 by market researcher TrendForce said Nvidia had moved beyond its original 12-Hi HBM4E design and started evaluating four separate memory setups, including one that would cut the chip’s on-package memory by roughly a third. The shift, later detailed by Korean outlet Chosun Ilbo citing reporting from The Information, shows how tight the HBM market has become even for the company that buys more of it than anyone else. Nvidia has not finalized a specification and has not commented publicly on the reported changes.
The story matters beyond one product line. Rubin Ultra sits at the center of Nvidia’s 2027 data center roadmap, and a memory cut at that scale would ripple through cloud pricing, AI training budgets, and the broader hardware supply chain that already strained under this year’s HBM crunch. Here is what’s confirmed, what’s still speculative, and what it means for the AI buildout.
Four Configurations, One Problem: Not Enough Memory
According to TrendForce, Nvidia’s original plan called for Rubin Ultra to ship with 12-Hi HBM4E stacks, the densest and newest memory format available. Starting in the third quarter of 2026, the company began testing three alternatives alongside that original design: 8-Hi HBM4E, 12-Hi HBM4, and 8-Hi HBM4. TrendForce said the final specification had not been determined at the time of its report, and no configuration had been locked for production.
TrendForce attributed the reassessment to two overlapping problems. First, overall DRAM supply is expected to stay tight through 2027, squeezing the wafer capacity that both conventional memory and HBM draw from. Second, yield and validation work on 12-Hi HBM4E remains unresolved, adding risk to a design that was supposed to anchor Nvidia’s next flagship accelerator. Put together, those two constraints pushed Nvidia toward testing lower-density fallback options rather than waiting on a single high-risk memory format.
The 192GB Configuration: What a 33% Cut Actually Means
The most-cited configuration among the four uses 8-Hi HBM4 stacks and lands around 192GB of total memory. Research firm SemiAnalysis, cited by Chosun Ilbo on September 16, 2026, described this 8-layer HBM4/192GB setup as the primary configuration currently being shared with customers for evaluation. Other sampled variants reportedly range between 192GB and 256GB, according to Korean business outlet Sedaily, which cited sourcing from The Information.
Compare that with Nvidia’s current Rubin chip, which is already in mass production and covered in shattered.io’s earlier reporting on HBM4 yield gains powering the Rubin platform. Rubin ships with up to 288GB of HBM4 in most reporting, though some accounts put maximum configurations as high as 384GB using larger 48GB stacks. Measured against the 288GB figure, a 192GB Rubin Ultra would represent roughly a 33% reduction, a number repeated across multiple outlets including Sedaily and technology site CryptoBriefing.
The gap looks even wider against Nvidia’s original ambitions. Korean reporting says Rubin Ultra was initially pitched with roughly 1TB of memory when the platform was first outlined. If that early design target holds up, a 192GB shipping configuration would fall about 81% short of it. That comparison deserves a caveat: a design target floated early in a product’s development cycle is not the same as a locked specification, and Nvidia has offered no public confirmation of either figure.
Why the Memory Squeeze Is Hitting Now
High-bandwidth memory is produced by three companies worldwide: SK Hynix, Samsung, and Micron. All three have spent 2026 racing to convert wafer capacity toward HBM3E and HBM4 as AI accelerator demand outstrips what conventional planning cycles anticipated. Nvidia’s own success is part of what’s straining supply. Strong results for the current Rubin platform, including Vera Rubin’s MLPerf inference showing versus Blackwell-generation GB300 systems, have pulled forward customer orders and locked up allocation that might otherwise have cushioned the Rubin Ultra transition.
That demand pressure collides with a supply side that simply can’t expand fast enough. New fabrication capacity takes years to bring online, and HBM’s stacked, high-yield-sensitive manufacturing process is harder to scale than standard DRAM. TrendForce’s assessment that DRAM supply stays tight “through 2027” effectively rules out a quick fix. Nvidia’s four-configuration hedge on Rubin Ultra looks like a direct response to that timeline: rather than betting the whole product on a memory format that might not be ready in volume, the company is keeping cheaper, more available fallback options on the table.
HBM Prices Are Already Spiking
The scarcity shows up clearly in pricing data. A 36GB HBM3E module traded at roughly $2,100 on the spot market as of September 4, 2026, according to research site IntuitionLabs. That’s four to five times the $300 to $400 range typically associated with long-term supply agreements between memory makers and their largest customers. Spot pricing spikes like that usually signal that buyers who missed out on committed allocation are paying a steep premium just to secure any supply at all, a dynamic that tends to precede broader price increases across the market.
That gap between spot and contract pricing also explains why Nvidia would rather redesign a chip than absorb the cost. Locking Rubin Ultra to a memory format still fighting yield problems risks both delayed shipments and margin pressure if the company has to chase spot-market supply to hit volume targets.
China’s Chipmakers Are Feeling the Same Squeeze
Nvidia isn’t alone. Reuters reported on September 10, 2026 that Chinese AI-chip makers, including Huawei and Cambricon, have started raising prices as high-bandwidth memory becomes harder to source. Huawei’s upcoming Ascend 950DT accelerator has reportedly been quoted at more than 250,000 yuan, or roughly $37,255, marking an increase of about 20% to 50% over earlier pricing according to Reuters’ sourcing. Shattered.io covered the broader pricing fallout in its earlier report on China’s AI chip price hikes, and the Rubin Ultra reconfiguration adds a second data point showing the shortage cuts across export-control lines and company borders alike.
That’s a notable detail for an industry often framed around the split between US-led and Chinese AI hardware ecosystems. Both sides are drawing from the same three-company HBM supply base, and both are now adjusting product plans or prices in response to the same underlying scarcity.
Nvidia Already Signaled the Supply Crunch Was Coming
The Rubin Ultra memory story doesn’t arrive out of nowhere. Shattered.io previously reported on how Nvidia CEO Jensen Huang’s public forecast of doubling chip output tied up a large share of available HBM supply months before this latest reporting surfaced. Locking in that much future capacity is a rational move for a company that expects demand to keep climbing, but it also means less room to maneuver when a specific product, like Rubin Ultra, runs into its own memory-format problems. The two stories reinforce each other: Nvidia secured aggressive supply commitments at the platform level, then still had to redesign an individual chip’s memory configuration when the numbers didn’t line up.
Ripple Effects: Gaming GPUs Pushed Further Out
The memory crunch isn’t confined to data-center chips. Leaker Moore’s Law Is Dead, cited by Techspot, said Nvidia’s next-generation RTX 60 series and most of AMD’s RDNA 5-based Radeon 10000 cards have reportedly been pushed back toward 2028, a delay shattered.io covered when both companies’ gaming roadmaps slipped. The stated reasoning lines up with the Rubin Ultra story: AI accelerators generate far higher margins per wafer than consumer graphics cards, so foundries and memory suppliers naturally prioritize data-center orders when capacity runs short.
Gamers waiting on next-generation cards are effectively competing with hyperscale AI buyers for the same limited pool of advanced packaging capacity and memory. Until HBM and leading-edge DRAM supply loosens, that competition tends to resolve in favor of whichever product carries the higher price tag, and data-center accelerators win that comparison easily.
Wafer Capacity Is the Underlying Bottleneck
Beneath the HBM headlines sits a more basic constraint: there simply isn’t enough advanced wafer capacity to satisfy conventional DRAM and HBM demand at the same time. TrendForce’s own framing ties the Rubin Ultra reassessment directly to that dynamic, describing DRAM supply broadly, not just HBM specifically, as tight through 2027. Shattered.io’s earlier coverage of silicon wafer prices potentially surging in 2027 on AI demand flagged the same underlying pressure from a different angle. Every HBM stack Samsung, SK Hynix, or Micron builds uses wafer capacity that could otherwise produce standard DRAM for phones, laptops, and servers, which is part of why memory costs have climbed well beyond the AI accelerator market this year.
Historical Context: Memory Shortages Aren’t New, But This One’s Different
The semiconductor industry has lived through supply crunches before. The 2017-2018 period saw a sustained DRAM and NAND pricing cycle driven by underinvestment relative to demand. The 2021 chip shortage hit a much broader swath of the industry, disrupting automotive, consumer electronics, and industrial production simultaneously. Both episodes eventually resolved as manufacturers added capacity and demand growth slowed.
This shortage is narrower in scope but arguably sharper in consequence. HBM isn’t a commodity component that can be substituted easily. It’s a purpose-built, high-margin product tied to a small number of advanced accelerators, and it sits directly on the critical path for the AI infrastructure buildout that dominates tech-sector capital spending. A shortfall here doesn’t just delay a product launch. It reshapes what the flagship chip of the next generation actually looks like, which is exactly what appears to be happening to Rubin Ultra.
Competitive Comparison: Rubin vs. Rubin Ultra Memory Configurations
The table below lays out the reported configurations side by side, distinguishing between what’s confirmed shipping today, what was originally planned, and what’s currently under evaluation.
| Chip / Configuration | HBM Generation | Stack Height | Reported Memory | Status |
|---|---|---|---|---|
| Rubin (current) | HBM4 | 8-Hi | Up to 288GB (some reports cite up to 384GB) | Shipping in 2026 |
| Rubin Ultra (original plan) | HBM4E | 12-Hi | ~1TB (early design target) | Superseded, per Korean reporting |
| Rubin Ultra (test config) | HBM4E | 8-Hi | Under evaluation | One of four options, per TrendForce |
| Rubin Ultra (test config) | HBM4 | 12-Hi | Under evaluation | One of four options, per TrendForce |
| Rubin Ultra (primary shared config) | HBM4 | 8-Hi | ~192GB (33% below Rubin’s 288GB) | Reported by SemiAnalysis as customer-facing baseline |
Nvidia’s rivals face a version of the same math. AMD’s Instinct accelerator line also depends on HBM4 allocation from the same three suppliers, and public reporting hasn’t yet detailed whether AMD has made comparable configuration changes for its own upcoming chips. What is clear is that no major AI accelerator maker sources HBM outside the SK Hynix, Samsung, and Micron trio, which means every company building flagship AI silicon in 2027 is negotiating around the same finite pool of stacked memory.
HBM Supply and Pricing Snapshot
| Metric | Figure | Source / Date |
|---|---|---|
| 36GB HBM3E spot price | ~$2,100 | IntuitionLabs, Sept. 4, 2026 |
| Typical long-term contract price (36GB HBM3E) | $300–$400 | Long-term supply agreements |
| Spot vs. contract price ratio | 4x–5x | IntuitionLabs analysis |
| Huawei Ascend 950DT quoted price | >250,000 yuan (~$37,255) | Reuters, Sept. 10, 2026 |
| Ascend 950DT price increase | ~20%–50% | Reuters reporting |
| DRAM supply outlook | Tight through 2027 | TrendForce, Aug. 4, 2026 |
| Number of global HBM suppliers | 3 (SK Hynix, Samsung, Micron) | Market structure, multiple outlets |
Market Impact: What This Means for AI Cloud Costs
Less memory per accelerator has direct consequences for anyone renting AI compute. Large language models are typically split across multiple GPUs when they can’t fit in one chip’s memory pool, and that sharding adds networking overhead and reduces efficiency. A Rubin Ultra shipping with 192GB instead of a hoped-for 1TB, or even instead of Rubin’s own 288GB, would push some workloads that might have fit on fewer chips back onto larger multi-GPU clusters. That tends to raise the effective cost of serving big models, even if the sticker price per chip stays flat.
Cloud providers already absorbed one round of memory-driven price increases this year. A leaner Rubin Ultra memory footprint gives them less room to offer the generational cost-per-token improvements that typically justify upgrading a data center’s accelerator fleet, which could stretch out the replacement cycle for current-generation Rubin and Blackwell-class hardware well into 2027 and 2028.
What Industry Trackers Are Watching Next
TrendForce’s initial report set the baseline: four configurations, no final decision, and a supply outlook that doesn’t clear up before 2027. SemiAnalysis has since narrowed the picture by identifying the 8-Hi HBM4, 192GB setup as the version most frequently shared with customers for evaluation, which suggests it may be the front-runner even without formal confirmation. The Information’s sourcing, relayed through Chosun Ilbo and Sedaily, adds the detail that some tested samples run as low as 192GB and as high as 256GB, meaning Nvidia is still testing a range rather than a single fallback number.
None of these outlets have reported a finalized specification, a shipping date narrower than “2027,” or an on-record comment from Nvidia. That leaves real uncertainty about where Rubin Ultra lands, and any of the four tracked configurations remains possible until Nvidia confirms otherwise.
Predictions: Where Rubin Ultra Goes From Here
- The 192GB, 8-Hi HBM4 configuration likely ships as the default SKU. It’s the version SemiAnalysis says is already circulating with customers, and it uses the more mature, higher-yield HBM4 format rather than the still-unproven 12-Hi HBM4E.
- Nvidia will probably segment the Rubin Ultra lineup rather than pick one configuration for everyone. Expect a higher-memory HBM4E variant reserved for premium customers once yields improve, alongside a HBM4-based mainstream version.
- HBM4E 12-Hi production issues likely persist into at least mid-2027. TrendForce’s “tight through 2027” framing suggests the format Nvidia originally wanted won’t be broadly available on the timeline the company first planned.
- Gaming GPU delays toward 2028 will likely hold or extend further. As long as data-center accelerators pay more per wafer than consumer cards, foundries and memory suppliers have little incentive to reprioritize.
- Expect further price increases across the AI accelerator market, not just at Nvidia. With Huawei and Cambricon already raising prices and all HBM sourced from the same three suppliers, competitors without Nvidia’s negotiating leverage may face even sharper cost pressure.
Frequently Asked Questions
Has Nvidia confirmed the Rubin Ultra memory cut?
No. The reporting comes from TrendForce, SemiAnalysis, and Korean outlets citing The Information. Nvidia has not made a public statement confirming any specific memory configuration for Rubin Ultra, and the company has not announced a final specification.
How much memory does the current Rubin chip have?
Most reporting puts Rubin, which is already in mass production in 2026, at up to 288GB of HBM4, though some accounts describe configurations reaching 384GB with larger memory stacks.
Why is Nvidia considering less memory instead of just waiting for more supply?
TrendForce’s reporting points to two factors: DRAM supply is expected to stay tight through 2027, and the 12-Hi HBM4E format Nvidia originally wanted still has unresolved yield and validation issues. Testing lower-density fallback configurations lets Nvidia keep Rubin Ultra on schedule rather than tying the launch to a memory format that isn’t ready.
When is Rubin Ultra expected to launch?
Reporting places Rubin Ultra in Nvidia’s 2027 roadmap, but no outlet has published a specific launch or shipping date. The final memory configuration and timing both remain unconfirmed.
Is this shortage affecting companies besides Nvidia?
Yes. Reuters reported that Chinese chipmakers Huawei and Cambricon have raised prices on their own AI accelerators as HBM has become harder to source, showing the shortage extends across the industry rather than being specific to one company or region.
Does this affect gaming GPU availability too?
Indirectly, yes. Reports citing leaker Moore’s Law Is Dead say Nvidia’s RTX 60 series and AMD’s RDNA 5-based Radeon 10000 cards have been pushed toward 2028 as foundries and memory suppliers prioritize higher-margin AI accelerator orders over consumer graphics cards.
Who actually makes HBM memory?
Three companies produce essentially all of the world’s high-bandwidth memory: SK Hynix, Samsung, and Micron. Every major AI accelerator maker, including Nvidia and AMD, sources HBM from that same small group of suppliers.
Could Nvidia still ship the original 1TB, 12-Hi HBM4E design?
It hasn’t been ruled out. TrendForce describes four configurations still under evaluation, including the original 12-Hi HBM4E plan. But given the reported yield and supply issues tied to that format, industry sourcing suggests a lower-memory HBM4-based configuration is the more likely outcome for the primary shipping version.




