Nvidia’s next data-center platform is no longer a slide at a keynote. Vera Rubin, the chip family built to succeed Blackwell, has moved into full production, and the first NVL72 racks are already running inside cloud data centers rather than test labs. Nvidia’s own newsroom put it plainly: the company said Vera Rubin “is ramping into full production, with Taiwan’s top server makers and global supply chain leaders manufacturing Vera Rubin-based systems at scale — fueling AI labs, cloud providers and hyperscalers to build tomorrow’s intelligence.” That single sentence, published on Nvidia’s data-center product page, is the clearest signal yet that the AI infrastructure race is entering its next round on schedule.

For a company that shipped a record $96.2 billion quarter earlier this year (as covered in our report on Nvidia’s Q2 earnings), the timing matters. Blackwell and its GB300 Ultra refresh are barely a year into volume shipments, and Nvidia is already pushing the next platform into hyperscaler racks. This piece breaks down what Vera Rubin NVL72 actually is, what’s confirmed versus what’s still marketing language, who is getting the first racks, and what it means for the rest of the AI hardware market heading into 2027.

What Nvidia Vera Rubin NVL72 actually is

Vera Rubin is Nvidia’s codename for the platform generation that follows Blackwell (B200/GB200) and the Blackwell Ultra refresh (GB300). It’s named, as with prior generations, after a scientist — in this case astronomer Vera Rubin, who provided some of the first strong evidence for dark matter. Where Blackwell paired Nvidia’s GPU with a Grace CPU, Vera Rubin pairs a new Rubin GPU with a new Vera CPU, and the whole thing ships as a rack-scale system rather than a single chip you can buy off a shelf.

Nvidia first showed Rubin publicly at GTC 2025, then used its CES 2026 keynote to confirm the platform had entered production. Jensen Huang told the CES audience the platform had moved past sampling and qualification straight into full-scale manufacturing, a jump some coverage described as faster than the industry expected for a chip this size. By March 2026, at GTC, Nvidia’s newsroom announced that “NVIDIA Vera Rubin platform is opening the next frontier of agentic AI, with seven new chips now in full production to scale the world’s largest AI factories.” Those seven chips reportedly span the Rubin GPU itself, the Vera CPU, networking silicon, and the NVSwitch 6 fabric that ties a rack together — a shift Nvidia frames as selling a complete system rather than a single part.

The NVL72 rack: what’s inside a single cabinet

The core product people are searching for is the Vera Rubin NVL72 rack, and the spec sheet is the reason it’s generating attention. According to Nvidia’s own product page and corroborating breakdowns from The Register and Hashrate Index, a single NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs into one liquid-cooled cabinet, tied together by nine NVSwitch 6 blades across 18 compute blades. The Register’s CES coverage reported that the fabric delivers 3.6 TB/s of bandwidth to each GPU, describing it as twice the per-GPU interconnect bandwidth of the prior-generation rack platform.

Memory is the other headline number. Multiple sources, including The Register’s CES coverage and Nvidia-adjacent systems documentation, converge on roughly 20.7 TB of HBM4 memory pooled across the rack, delivering around 1,580 TB/s of aggregate HBM bandwidth. The Register reported that each Rubin GPU carries 288 GB of HBM4 — the same per-socket capacity as the outgoing Blackwell Ultra (GB300) chip, but running at roughly 2.8 times the bandwidth, at 22 TB/s per socket. On top of the HBM4 pool, the rack adds 54 TB of LPDDR5X spread across its compute blades for additional capacity.

On raw throughput, industry breakdowns citing Nvidia’s own figures put a single NVL72 rack at roughly 3,600 PFLOPS of NVFP4 inference performance and 2,520 PFLOPS of NVFP4 training performance — figures that, if accurate, would put one rack well into exaFLOPS territory for the low-precision math formats that dominate large language model inference today. Nvidia has not published a rack-level power draw for NVL72 in the sources reviewed for this piece, so any power figure circulating online should be treated as unverified until Nvidia or a named reviewer confirms it directly.

Vera Rubin NVL72 vs Blackwell GB300 NVL72

The generational jump is the part buyers actually care about, since most hyperscalers are choosing between finishing out their Blackwell orders or waiting for Rubin allocation. Comparative specs published by systems integrator Delta Computer Products show the GB300 NVL72 rack carrying roughly 21 TB of HBM3e memory and delivering about 1.1 exaFLOPS of FP4 performance. Against that baseline, Vera Rubin NVL72’s reported 3.6 exaFLOPS of NVFP4 performance represents more than three times the FP4-class throughput of the outgoing Blackwell Ultra rack, while trading HBM3e for faster HBM4 at a similar total capacity.

SpecBlackwell Ultra (GB300 NVL72)Vera Rubin NVL72
GPUs per rack7272
CPU pairingGrace (36 CPUs)Vera (36 CPUs)
Memory type~21 TB HBM3e~20.7 TB HBM4
Per-socket GPU bandwidthBaseline (prior generation)~22 TB/s (~2.8x per The Register)
NVFP4 performance (rack)~1.1 EFLOPS FP4~3.6 EFLOPS NVFP4
Additional system memoryNot specified in sources reviewed54 TB LPDDR5X
Production status (Sept. 2026)Shipping since 2025Full production, partner shipments underway

It’s worth noting none of the sources reviewed for this piece put a direct, Nvidia-confirmed FLOPS comparison against the older Hopper H100 generation — comparisons in Nvidia’s own materials and in outlets like The Register are framed against Blackwell, not Hopper. Readers comparing Rubin to H100 directly should treat multi-generation leaps cited elsewhere as industry estimates rather than confirmed Nvidia figures, since Hopper-to-Rubin numbers weren’t part of the on-record material for this article.

Timeline: how fast Nvidia actually moved

The production timeline is unusually compressed even by Nvidia’s own annual cadence. Systems integrator Introl, tracking Nvidia’s public statements, laid out the milestones as production qualification in the fourth quarter of 2025, full production starting in the first quarter of 2026, and cloud availability arriving in the second half of 2026. Nvidia’s own January 5, 2026 newsroom statement backs that framing directly: “NVIDIA Rubin is in full production, and Rubin-based products will be available from partners the second half of 2026.”

From there, the dates tighten further. Nvidia’s March 2026 GTC announcement confirmed seven Rubin-platform chips in full production. Newsletter AI Weekly, citing manufacturing data, reported in June 2026 that Vera Rubin production had scaled across more than 350 factories in 30 countries, backed by more than 150 Taiwan-based ecosystem partners, with shipments beginning in fall 2026 through OEMs including Dell Technologies, HPE, Lenovo, Supermicro, ASUS, Foxconn, and Gigabyte. By July 22, 2026, Korean outlet ChosunBiz reported that Nvidia had begun full-scale deliveries, stating the company “supplied the next-generation AI rack ‘Vera Rubin NVL72’ to Google Cloud, Microsoft Azure, Oracle Cloud, and CoreWeave,” with those systems already running at customer sites.

DateMilestoneSource
GTC 2025Vera Rubin platform first shown publiclyNvidia newsroom
Q4 2025Production qualification (per tracked timeline)Introl
Jan. 5, 2026Huang confirms full-scale production at CES keynoteNvidia newsroom / Yahoo Tech
March 16, 2026Seven Vera Rubin chips confirmed in full production at GTCNvidia newsroom
June 1, 2026Production scaled across 350+ factories, 30 countriesAI Weekly
July 21-22, 2026First NVL72 racks delivered to Google Cloud, Azure, Oracle, CoreWeaveChosunBiz
Aug. 7, 2026Nvidia confirms ramp into full production on official product pageNvidia.com
Fall 2026Broader OEM channel shipments beginAI Weekly, Silicon Report

The dates aren’t perfectly consistent across every outlet — some pin “full production” to Q1 2026, others to June 1, 2026 — but every source reviewed agrees on the broader shape: qualification in late 2025, production ramp through the first half of 2026, and named hyperscaler deliveries by midsummer. That’s a tighter gap between platform reveal and hyperscaler deployment than Blackwell managed, and it’s happening while Blackwell Ultra racks are still shipping in volume.

Who’s getting the first racks

ChosunBiz named four hyperscalers as recipients of the first Vera Rubin NVL72 systems as of late July 2026: Google Cloud, Microsoft Azure, Oracle Cloud, and CoreWeave. A separate breakdown from Silicon Report expanded that list to include AWS, Lambda, Nebius, and Nscale as partners slated for shipments beginning in fall 2026, though that report frames those additional names as forward-looking allocation rather than confirmed deliveries already in production. Notably, none of the reviewed sources listed Meta, xAI, or Anthropic by name as direct Vera Rubin recipients — those companies’ compute strategies for this generation weren’t part of the on-record material gathered for this piece.

On the manufacturing side, the supply chain looks similar to Blackwell’s, just larger. Silicon Report and a Robotics Media report both name Foxconn, Quanta, and Wistron as the Taiwanese ODM partners scaling rack-level production, while TSMC is reported to be mass-producing the Rubin dies on its 3nm node with capacity booked through 2027. On memory, SK hynix is named as a supplier shipping 12-layer HBM4E samples intended for later Rubin variants — though the sources reviewed did not name Samsung or Micron directly in connection with the current NVL72 memory supply, a gap worth watching given both companies compete aggressively for HBM allocation, as we covered when HBM4 yields hit 80% for Nvidia’s Rubin ramp and when Samsung built 8-layer HBM4E for Nvidia earlier this year.

Why this matters for the AI infrastructure market

Nvidia’s data-center business has become the single biggest driver of the company’s revenue, and the cadence of new platforms is now central to how fast that revenue can keep growing. The company doesn’t appear to be waiting for Blackwell demand to plateau before pushing the next platform into hyperscaler racks — a strategy that keeps competitors chasing a moving target. It also puts pressure on the rest of the supply chain: server makers, memory suppliers, and cooling vendors all have to keep pace with a company shipping new rack architectures roughly once a year rather than once every two to three years, the cadence that was typical earlier in the GPU era.

That pace also raises the stakes for anyone who bought into the previous generation. Cloud providers that locked in large GB300 orders now have to decide how much additional capacity to commit to Rubin racks that promise more than triple the FP4 throughput. Server vendors named in the production ramp, including HPE, Dell, Lenovo, and Supermicro, have a direct incentive to get Rubin-based SKUs to market fast; we’ve already tracked how HPE stock hit a 52-week high on its Nvidia Vera CPU server plans, a sign that investors are pricing in Rubin-cycle revenue well before broad retail availability.

There’s a supply-side risk baked into all of this too. Nvidia has already had to raise prices on some AI server components this year, citing a broader memory shortage, as detailed in our coverage of Nvidia’s 15%+ AI server price hikes tied to the memory crunch. A rack architecture that needs 20.7 TB of leading-edge HBM4 per unit, multiplied across hundreds of racks per hyperscaler order, puts even more strain on a memory market that was already tight before Rubin ramped.

Competitive pressure: AMD, Google, and custom silicon

Vera Rubin’s ramp lands at a moment when Nvidia’s biggest customers are simultaneously trying to reduce their reliance on Nvidia. As we reported when covering how Nvidia’s biggest customers are now building five rival chip programs, Google, Amazon, Microsoft, and Meta have all pushed forward their own custom AI accelerators over the past year, aiming to cut both cost and Nvidia’s leverage over their infrastructure roadmaps. AMD, meanwhile, continues to compete on the merchant-silicon side with its Instinct accelerator line, positioning itself as the primary alternative for buyers who don’t want to build custom chips but also don’t want to pay full price for Nvidia’s newest rack.

None of that competitive pressure appears to have slowed Vera Rubin’s rollout. If anything, the compressed timeline from GTC 2025 reveal to named hyperscaler deployments by July 2026 suggests Nvidia is trying to out-execute rivals on release cadence rather than compete purely on price. Whether that strategy holds depends heavily on whether TSMC’s 3nm capacity and the broader HBM4 supply chain can keep up with demand that, per AI Weekly’s reporting, is already being fed by more than 350 factories worldwide.

Historical context: from Hopper to Rubin in three generations

It’s easy to lose track of how fast Nvidia’s data-center architecture has turned over. Hopper (H100/H200) launched in 2022 and became the default AI training chip through 2023 and into 2024. Blackwell (B200/GB200) followed in 2024, doubling down on rack-scale NVLink systems rather than individual accelerator cards. Blackwell Ultra (GB300) arrived as a mid-cycle refresh in 2025, adding more memory and modest performance gains without a full architectural overhaul. Vera Rubin, entering full production in 2026, is the first generation since Hopper to pair a genuinely new GPU architecture with a new custom CPU (Vera, succeeding Grace) and a new switch generation (NVSwitch 6) all at once.

That three-chip refresh, shipping together as a system rather than staggered component upgrades, is part of why Nvidia frames Rubin as a platform rather than a product. It also explains the company’s emphasis on manufacturing scale in its own statements — moving 20.7 TB of HBM4 memory, 72 GPUs, and 36 CPUs through validation and into a single rack is a materially harder manufacturing problem than shipping a single new GPU die.

What Nvidia and industry voices are saying

Nvidia’s public messaging around Vera Rubin has stayed consistent from CES through its August product-page update. The company’s official announcement stated: “NVIDIA Rubin is in full production, and Rubin-based products will be available from partners the second half of 2026” (Nvidia Newsroom). Its later data-center product page reiterated that Vera Rubin “is ramping into full production, with Taiwan’s top server makers and global supply chain leaders manufacturing Vera Rubin-based systems at scale — fueling AI labs, cloud providers and hyperscalers to build tomorrow’s intelligence” (Nvidia Newsroom).

CEO Jensen Huang framed the platform’s purpose during keynote remarks covered by SiliconANGLE, saying Vera Rubin “was built for this moment – an AI factory engine that delivers intelligence at scale, with the performance, efficiency and security needed to power the next industrial revolution” (SiliconANGLE). Huang also tied the platform’s design to the demands of autonomous, agentic AI workloads, noting that “agentic AI turns enterprise data into a living, real-time system – and that system must be protected where data moves, where context is stored and where agents act” (SiliconANGLE). Taken together, the two remarks show Nvidia positioning Rubin less as a raw-performance upgrade and more as infrastructure purpose-built for the agentic AI workloads it expects to dominate enterprise spending through 2027.

What’s confirmed and what’s still unverified

Given how much speculative hardware coverage circulates around any Nvidia launch, it’s worth separating what’s on the record from what isn’t. Confirmed by Nvidia directly: the platform is in full production, seven chips are part of the Vera Rubin family, and shipments to partners are underway in the second half of 2026. Confirmed by named third-party reporting: specific hyperscaler recipients (Google Cloud, Azure, Oracle, CoreWeave), the 72-GPU/36-CPU rack configuration, and the approximate HBM4 and LPDDR5X capacity figures.

Not confirmed in the sources reviewed for this article: an official per-rack power consumption figure, official pricing or per-unit revenue numbers, and a direct Nvidia-published performance comparison against Hopper H100. Readers should treat any specific dollar price or kilowatt power figure they see elsewhere for Vera Rubin NVL72 as unverified until Nvidia publishes it directly, since none of the outlets checked for this piece — including Nvidia’s own newsroom and product pages — included those numbers as of publication.

Predictions: what happens next

  • Expect Nvidia to publish official pricing guidance or at least ASP commentary during its next quarterly earnings call, since analysts will push for Rubin’s contribution to guidance now that named hyperscaler shipments are underway.
  • Watch for additional hyperscaler names beyond the four confirmed by ChosunBiz (Google Cloud, Azure, Oracle, CoreWeave) to be confirmed as recipients before the end of 2026, given Silicon Report’s reporting that AWS, Lambda, Nebius, and Nscale are already in the allocation queue.
  • HBM4 supply will likely remain the primary bottleneck through early 2027, with SK hynix’s sampling of 12-layer HBM4E suggesting even denser memory stacks are coming for later Rubin variants rather than the initial NVL72 configuration.
  • Server OEMs named in the production ramp — Dell, HPE, Lenovo, Supermicro, ASUS, and Gigabyte — are likely to announce Rubin-based SKUs at their respective product events through the rest of 2026, following the pattern set by Blackwell’s rollout.
  • Competing custom-silicon programs from Google, Amazon, and Meta will likely accelerate in response, since none of those companies want to be locked into Rubin-cycle pricing for their next generation of internal AI training capacity.

What buyers and infrastructure teams should watch

For infrastructure teams evaluating whether to wait for Rubin allocation or take delivery of remaining Blackwell Ultra capacity, the practical calculus comes down to lead time and workload mix. Nvidia’s reported 2.8x per-socket bandwidth improvement over Blackwell Ultra will matter most for training workloads that are memory-bandwidth bound, while inference-heavy deployments may see a smaller relative benefit until software stacks fully exploit the new NVFP4 pipeline. Given that only four hyperscalers have confirmed deliveries as of this writing, most enterprise buyers renting capacity through cloud providers won’t have direct access to Rubin-based instances until those providers finish standing up their own NVL72 fleets, likely sometime in the fourth quarter of 2026 at the earliest based on the reported fall 2026 OEM shipment window.

Teams currently negotiating multi-year GPU capacity contracts should also factor in Nvidia’s roughly one-year cadence between major platform generations. Locking into a long-term Blackwell-only contract now, with Rubin racks already running at four major clouds, risks leaving a gap between contract renewal and next-generation access. That’s a different negotiating position than infrastructure buyers had during the Hopper era, when platform refreshes were spaced two to three years apart rather than one.

Frequently asked questions

What is Nvidia Vera Rubin NVL72?

It’s Nvidia’s next data-center rack platform, succeeding Blackwell and Blackwell Ultra (GB300). A single NVL72 rack combines 72 Rubin GPUs with 36 Vera CPUs, connected by NVSwitch 6 fabric, and is built around roughly 20.7 TB of HBM4 memory.

When did Vera Rubin enter full production?

Nvidia confirmed full production at its CES 2026 keynote in January, with its March 2026 GTC announcement confirming seven platform chips in full production. Different trackers cite dates ranging from Q1 2026 to June 1, 2026 for the full-production milestone, but Nvidia’s own newsroom and product pages confirm the platform is in production as of mid-2026.

Which companies have received Vera Rubin NVL72 racks first?

ChosunBiz reported that Google Cloud, Microsoft Azure, Oracle Cloud, and CoreWeave received the first confirmed NVL72 systems as of late July 2026. Other outlets report AWS, Lambda, Nebius, and Nscale as additional partners slated for shipments beginning in fall 2026, though those weren’t confirmed as already-delivered at the time of that reporting.

How does Vera Rubin NVL72 compare to Blackwell GB300 NVL72?

Per specs compiled by systems integrators, Vera Rubin NVL72 delivers roughly 3.6 exaFLOPS of NVFP4 performance versus about 1.1 exaFLOPS of FP4 performance on GB300 NVL72 — more than triple the throughput — while moving from HBM3e to faster HBM4 memory at a similar total capacity.

How much HBM4 memory does a Vera Rubin NVL72 rack use?

Reported figures converge on roughly 20.7 TB of HBM4 memory per rack, delivering around 1,580 TB/s of aggregate bandwidth, plus an additional 54 TB of LPDDR5X spread across the rack’s compute blades.

Is Vera Rubin NVL72 pricing available yet?

No official per-rack or per-unit price has been published in the sources reviewed for this article. Any specific dollar figures circulating for Vera Rubin NVL72 should be treated as unverified estimates rather than confirmed Nvidia pricing.

Who is manufacturing Vera Rubin NVL72 systems?

TSMC is reported to be manufacturing the Rubin dies on its 3nm process, with Foxconn, Quanta, and Wistron named as the Taiwanese ODM partners assembling full racks. SK hynix is named as a memory partner supplying HBM4E samples for later Rubin variants. Server makers named in the broader production ramp include Dell Technologies, HPE, Lenovo, Supermicro, ASUS, and Gigabyte.

Does Vera Rubin replace Blackwell immediately?

No. Blackwell and Blackwell Ultra racks continue shipping in volume alongside the Vera Rubin ramp. Most cloud capacity available to enterprise customers today still runs on Blackwell-generation hardware, since Vera Rubin deployments so far are concentrated among a small number of named hyperscalers building out their own fleets.