Nvidia said on May 31, 2026 that its next-generation Vera Rubin AI computing platform is “ramping into full production,” with the first customer shipments set to begin this fall. The announcement, posted on Nvidia’s newsroom site, marks the moment the company’s most ambitious data-center platform to date moves from lab demos and reference designs into racks that will actually run inside cloud data centers. For a chip industry that spent 2025 absorbing Blackwell, the shift to Rubin comes fast, and it arrives with a long list of named customers already attached.

The stakes go beyond one product launch. Nvidia built its position atop Blackwell’s dominance in AI training and inference, but 2026 has brought sharper competition from AMD’s Helios rack, Google’s in-house TPUs, and a supply chain still working through HBM4 memory constraints. How fast Rubin actually ships, and to whom, will shape AI hardware spending for the next two years. Here is what is confirmed, what remains a vendor claim, and what to watch next.

What “Full Production” Actually Means for Nvidia Vera Rubin

“Full production” is a specific manufacturing milestone, not a marketing phrase. It means Nvidia and its assembly partners are building Vera Rubin systems at scale rather than producing small batches of engineering samples. Nvidia’s own newsroom post states plainly that the platform is “ramping into full production to power agentic AI factories worldwide,” and that production shipments are set to begin this fall, which in Nvidia’s fiscal calendar lines up with the third quarter.

That timeline tracks with earlier checkpoints. Nvidia gave a platform deep dive in January 2026 that outlined a six-to-seven chip architecture built around the NVL72 rack. By March, at GTC, the company said new Rubin-era chips were already entering full production to scale what it calls the world’s largest AI factories. The May 31 announcement is the point where that manufacturing status became a formal, dated commitment tied to specific system builders and cloud customers rather than a roadmap slide.

Nvidia CEO Jensen Huang framed the platform around a specific workload category: agentic AI, where models plan multi-step tasks and call tools rather than simply answering a single prompt. That framing matters for how Nvidia is pricing and positioning Rubin against Blackwell, since agentic workloads tend to push more tokens through a system per user session than a single chatbot reply does.

Inside the Platform: Vera CPU, Rubin GPU, and Five-Rack Supercomputers

Vera Rubin is not a single chip. It is a platform built around two custom parts, a Vera CPU and a Rubin GPU, paired inside a rack-scale system Nvidia calls NVL72. Each Rubin GPU carries 288GB of HBM4 memory with roughly 22 terabytes per second of memory bandwidth per GPU, connected through NVLink 6 at 3.6 terabytes per second per GPU. A full NVL72 rack packs 72 Rubin GPUs alongside their paired Vera CPUs, for a combined 20.7 terabytes of pooled HBM4 memory and a rated 3,600 petaflops, or 3.6 exaflops, of NVFP4 inference throughput per rack.

Nvidia’s newsroom post describes the broader deployment unit as five purpose-built racks working together as one unified AI supercomputer, not just a single NVL72 cabinet. The company also highlighted its Spectrum-X Ethernet Photonics networking gear as part of the package, claiming five times better power efficiency, five times longer uptime, and 1.3 times faster deployment compared with traditional optical transceivers. None of those photonics figures have been independently benchmarked outside Nvidia’s own materials, so they should be read as vendor claims rather than confirmed third-party results.

The headline performance number Nvidia is pushing hardest is a 10x gain in agent throughput compared with the Grace Blackwell platform it replaces. Cloud provider CoreWeave, cited in coverage of the rollout, separately reported roughly ten times the token output from NVL72 racks compared with the previous generation, a figure that lines up with Nvidia’s own claim but again originates from a customer briefing rather than an independent lab test.

The HBM4 Scramble: Three Memory Makers, One Deadline

Rubin’s memory-heavy design put Nvidia in an unusual position for 2026: it needed all three major HBM suppliers running at once just to keep pace with its own shipment schedule. Speaking to reporters in June, Huang confirmed that SK Hynix, Samsung, and Micron had each been qualified to supply HBM4 for Vera Rubin. “All three vendors have been qualified,” Huang said, adding that “all three vendors are in production, and they’re all racing to support Vera Rubin,” according to a report from Tech Times.

That three-way qualification is a departure from the Blackwell era, when Nvidia leaned more heavily on a narrower supplier base and periodically ran into HBM3e bottlenecks that slowed GB200 deliveries through parts of 2025. Spreading HBM4 sourcing across SK Hynix, Samsung, and Micron simultaneously gives Nvidia more room to hit its fall shipment target even if one supplier’s yields lag, though it also means Rubin’s actual output depends on three separate manufacturing ramps staying roughly in sync. A companion piece on HBM4 yield rates feeding the Rubin ramp has more detail on how each supplier’s production numbers compare.

Who Gets Chips First: The Cloud and OEM Customer List

Nvidia’s May 31 release names an unusually long list of partners for a hardware announcement. On the systems side, Dell Technologies, HPE, Lenovo, and Supermicro are listed as lead builders, alongside a broader group that includes Foxconn, ASUS, MSI, GIGABYTE, Wistron, Wiwynn, Pegatron, Inventec, Compal, and roughly a dozen more manufacturers and storage partners such as NetApp, VAST Data, and WEKA.

On the cloud side, Nvidia names CoreWeave, Microsoft Azure, Lambda, Nebius, Nscale, IBM Cloud, GMI Cloud, IREN, Firmus, SpaceXAI, and Vultr as adopters of the platform’s confidential-computing features specifically. Separate reporting on the rollout, including coverage from TNW, adds OpenAI, Google Cloud, Meta, Anthropic, Perplexity, SpaceX, and Oracle to the list of organizations expected to deploy Rubin systems, with OpenAI reportedly planning to adopt the platform at scale during the third quarter. That breadth of named customers, spanning both hyperscalers and newer GPU-cloud specialists, suggests Nvidia is trying to avoid the allocation bottlenecks that frustrated some Blackwell buyers in 2025, even if actual chip counts per customer have not been disclosed.

From Blackwell to Rubin: A Faster, Steeper Ramp

Nvidia unveiled Blackwell, including the B200 GPU and GB200 “Grace Blackwell” superchip, at GTC in March 2024. That generation moved through sampling and limited shipments over the following year before becoming the default high-end platform at most hyperscalers by sometime in 2025, a ramp that took roughly 18 months from stage announcement to broad hyperscaler availability. Nvidia’s own B200 spec sheet lists 192GB of HBM3e memory per GPU and roughly 8.0 terabytes per second of memory bandwidth, figures that became the baseline every 2025 AI server was measured against.

Rubin’s public timeline compresses that cycle. From the January 2026 platform deep dive to the May 31 full-production announcement is roughly five months, and Nvidia is targeting fall shipments within the same calendar year the platform was first detailed in depth. Whether that compressed schedule holds up once volume manufacturing actually starts is the open question, since Blackwell’s own ramp slipped against early guidance more than once in 2024 and 2025 as packaging and memory supply caught up with demand.

AMD’s Helios Rack Answers Back

Nvidia is not shipping into an empty market. AMD’s competing rack-scale system, Helios, built around the MI455X accelerator, is aimed squarely at the same hyperscaler buyers. Each MI455X carries 432GB of HBM4, and a full 72-GPU Helios rack totals roughly 31 terabytes of pooled memory, well above the Rubin NVL72’s 20.7TB. AMD has already shipped a related generation of the chip at scale: the company’s MI450 deal with Oracle put 50,000 GPUs into that cloud provider’s hands ahead of the Helios rack rollout.

AMD’s own materials claim Helios delivers up to 15% more AI compute, 50% more HBM capacity, and 50% more scale-out bandwidth than an Nvidia Vera Rubin NVL72 rack, with Microsoft named as a launch customer. That compute claim sits awkwardly next to Nvidia’s own published number of 3,600 petaflops of NVFP4 inference per Rubin NVL72 rack versus a reported 2.9 exaflops of FP4 dense compute for a Helios rack, a gap that likely reflects the two companies measuring different precision formats and workload mixes rather than a straightforward apples-to-apples result. What is clearer is the production calendar: Rubin racks began reaching hyperscalers over the summer, while Helios has largely remained at reference-design status through mid-2026, with AMD targeting broader shipments in the second half of the year.

Rubin vs. Helios vs. Blackwell: The Spec Comparison

Metric (per rack unless noted)Blackwell GB200 NVL72 (2024-25)Rubin NVL72 (2026)AMD Helios MI455X (2026)
GPUs per rack727272
Memory per GPU192GB HBM3e288GB HBM4432GB HBM4
Memory bandwidth per GPU8.0 TB/s22 TB/s19.6 TB/s
Total pooled rack memory~13.8 TB20.7 TB~31 TB
Peak rack inference computeNot directly comparable (per-GPU FP8 spec only)3,600 PFLOPS NVFP42.9 EF FP4 dense (AMD figure)
Full production status (mid-2026)Broadly shipped since 2025Confirmed May 31, 2026; ships fall 2026Reference-design stage; H2 2026 target

The table shows why the Rubin-versus-Helios comparison is not a clean win for either side. AMD leads on raw memory capacity per rack, which matters for serving very large mixture-of-experts models without splitting them across more hardware. Nvidia leads on time to market and on its own compute metric, and it enters the second half of 2026 with a longer, more diverse list of named cloud customers already attached to shipments.

Ironwood, Custom Silicon, and the Rest of the Field

Nvidia and AMD are not the only names building AI accelerators at rack scale in 2026. Google continues to push its own Ironwood TPU line for internal workloads and, increasingly, for external cloud customers renting TPU capacity directly against Nvidia GPU pricing. Microsoft has its own Maia accelerator program running in parallel with the Azure capacity it is buying from Nvidia, a dual-track strategy detailed in coverage of HPE’s stock move tied to Vera CPU-based servers. Several hyperscalers have also been reported to be exploring custom AI silicon designed with merchant chipmakers, though deal specifics for those programs remain unconfirmed as of this writing.

The practical effect for Nvidia is that Rubin has to win on more than raw specs. Cloud providers building their own silicon are, by definition, price-sensitive customers who will only keep buying Nvidia GPUs if the performance gap justifies the premium. Rubin’s 10x agent-throughput claim, if it holds up under third-party testing once systems ship in volume, is Nvidia’s argument for why that premium remains worth paying even as internal alternatives mature.

TSMC’s Bottleneck and the 150-Company Supply Chain

Every Rubin GPU still has to clear TSMC’s advanced packaging lines before it reaches a Dell or Supermicro chassis. Nvidia’s own materials describe Taiwan’s top server makers and “global supply chain leaders” manufacturing Rubin systems at scale, and a supply-chain breakdown from Introl names Foxconn, Quanta, and Wistron among roughly 150 companies involved in building Rubin racks. That same analysis estimates 2026 output may be capped somewhere in the 200,000-to-300,000 Rubin GPU range, a ceiling driven largely by TSMC’s N3 wafer capacity and advanced packaging throughput rather than by chip design constraints, according to the Introl supply-chain analysis.

That capacity figure is an outside estimate, not a number Nvidia has confirmed, but it is a useful reference point. If it holds, Rubin’s first-year volume would land well below the millions of GPU units Nvidia has shipped cumulatively across the Hopper and Blackwell generations combined, meaning early Rubin allocation will likely go first to the largest, longest-standing customers such as Microsoft, CoreWeave, and Google Cloud before smaller buyers see meaningful volume.

Vera Rubin Rollout Timeline

MilestoneDateDetail
Platform deep diveJanuary 2026Nvidia outlines the six-to-seven chip Rubin architecture and NVL72 rack design
GTC production updateMarch 2026Nvidia says new Rubin-generation chips are entering full production
HBM4 supplier confirmationJune 2026Jensen Huang confirms SK Hynix, Samsung, and Micron are all qualified and producing HBM4 for Rubin
Full production announcementMay 31, 2026Nvidia formally states Vera Rubin is “ramping into full production”
First hyperscaler shipmentsSummer 2026Early NVL72 systems begin reaching cloud partners including CoreWeave
Broad customer shipmentsFall 2026Nvidia says production shipments to the wider customer list begin this quarter
OEM server availabilitySecond half of 2026Dell, HPE, Lenovo, and Supermicro systems built on Vera Rubin reach the market

Market Reaction: Why Wall Street Keeps Watching Rubin

Rubin’s rollout lands at a moment when Nvidia’s data-center business is already the dominant driver of its results. The company posted a $96.2 billion quarterly revenue figure earlier in 2026 as AI infrastructure spending kept climbing, and investors have treated each new platform milestone as a signal of whether that growth rate can hold into 2027. AMD, for its part, used its own results this year to argue its addressable market has room to grow alongside Nvidia’s, not just at its expense, a case the company’s CFO made when it lifted its total addressable market estimate to $3 trillion.

The market’s real test for Rubin will not be the May 31 announcement itself but whether fall shipment volumes match what Nvidia has told customers to expect. Cloud GPU rental pricing, already volatile amid the transition from Hopper to Blackwell to Rubin, gives an early read on scarcity: elevated pricing for Rubin-class capacity once it appears on cloud marketplaces would confirm that demand is outrunning the roughly 150-company supply chain Nvidia has assembled to build it, while pricing closer to today’s B200 and H200 rental rates would suggest Nvidia has managed the ramp more smoothly than Blackwell’s own rollout.

What Nvidia and Its Executives Are Saying

Nvidia’s public statements around the Rubin ramp have stayed narrow and specific rather than sweeping. The company’s newsroom post states: “NVIDIA today announced the NVIDIA Vera Rubin platform is ramping into full production to power agentic AI factories worldwide,” according to the official Nvidia press release. The same release is direct about timing: “Production shipments of Vera Rubin are set to begin starting this fall.”

On the memory-supply question that has shadowed Rubin since early 2026, Huang has been similarly specific rather than promotional. Asked about HBM4 readiness, he told reporters “all three vendors have been qualified,” and followed up that “all three vendors are in production, and they’re all racing to support Vera Rubin,” as reported by Tech Times. A separate company filing summary reiterates the shipment language in fiscal-year terms: “Our next-generation Data Center architecture, Vera Rubin, began production shipments in the third quarter of fiscal year 2027,” according to a filing summary republished by whatledto.com. Taken together, the statements describe a company being careful to tie Rubin’s launch to a specific fiscal quarter rather than a vague “coming soon,” a shift from how Nvidia talked about early Blackwell delays in 2024.

Five Predictions for Rubin’s First Year on the Market

First, expect Nvidia’s named hyperscaler customers, particularly Microsoft Azure, CoreWeave, and Google Cloud, to absorb the bulk of 2026 Rubin allocation, leaving smaller GPU-cloud players to wait until 2027 for meaningful volume, consistent with how Blackwell allocation played out in its own first year.

Second, the gap between AMD’s Helios compute claims and Nvidia’s own Rubin NVFP4 numbers will likely narrow once both companies publish workload-matched, third-party benchmarks, rather than the vendor-sourced figures both sides are currently citing.

Third, HBM4 pricing and allocation will remain the tightest constraint on how fast Rubin can actually ship, even with three qualified suppliers, since packaging capacity at TSMC caps how many finished GPUs can reach customers regardless of memory availability.

Fourth, cloud rental pricing for Rubin-class capacity will likely launch at a premium over current B200 and H200 rates, then compress over two to three quarters as supply catches up, mirroring the pricing curve seen across the Hopper-to-Blackwell transition.

Fifth, expect at least one more named hyperscaler to disclose a custom-silicon program aimed at reducing Rubin dependence before the end of 2026, continuing the pattern already set by Google’s TPU line and Microsoft’s Maia chips.

What This Means for Buyers and Developers

For engineering teams planning infrastructure budgets, the practical takeaway is timing rather than specs. Rubin capacity is not broadly available yet, and Nvidia’s own language points to fall 2026 as the earliest realistic window for new customers outside the named launch partners. Teams currently running production workloads on Blackwell-generation B200 or H200 instances have no urgent reason to migrate before Rubin pricing and availability stabilize, particularly since early-generation pricing on new Nvidia platforms has historically carried a premium that fades within a few quarters.

Teams evaluating AMD’s Helios as an alternative should weigh its larger per-rack memory pool against its later production timeline, since a rack with more HBM4 capacity does not help if it is not shippable until the second half of 2026 at the earliest. For workloads that fit comfortably within existing Blackwell memory limits, the safest short-term move is to keep optimizing on current-generation hardware while watching how the Rubin-versus-Helios allocation picture develops over the next two quarters.

Frequently Asked Questions

What is Nvidia Vera Rubin?
Vera Rubin is Nvidia’s next-generation AI data-center platform, combining a custom Vera CPU with Rubin GPUs inside a rack-scale system called NVL72. It succeeds the Blackwell (B200/GB200) generation and is aimed primarily at agentic AI and large mixture-of-experts model workloads.

When does Vera Rubin ship?
Nvidia confirmed on May 31, 2026 that the platform is ramping into full production, with production shipments to its broader customer list set to begin in fall 2026. Some early systems reportedly reached cloud partners such as CoreWeave over the summer.

How much memory does a Rubin GPU have?
Each Rubin GPU carries 288GB of HBM4 memory with about 22 terabytes per second of bandwidth. A full 72-GPU NVL72 rack pools 20.7 terabytes of HBM4 across the system.

Who is supplying HBM4 memory for Vera Rubin?
Nvidia CEO Jensen Huang confirmed that all three major memory makers, SK Hynix, Samsung, and Micron, have been qualified and are in production to supply HBM4 for the platform.

How does Rubin compare to AMD’s Helios rack?
AMD’s Helios, built around the MI455X GPU, packs more memory per rack (roughly 31TB versus Rubin’s 20.7TB) and claims a compute and bandwidth edge in its own materials. Rubin currently leads on production timing, with shipments already underway to some cloud partners while Helios remains largely at reference-design stage through mid-2026.

Which companies are named as Vera Rubin customers?
Nvidia’s announcement names Dell, HPE, Lenovo, and Supermicro as lead system builders, and CoreWeave, Microsoft Azure, Lambda, Nebius, Nscale, and others as cloud adopters. Separate reporting adds OpenAI, Google Cloud, Meta, Anthropic, and Oracle to the list of organizations expected to deploy the platform.

Is the 10x agent throughput claim independently verified?
No. That figure comes from Nvidia’s own comparison against the Grace Blackwell platform, echoed by customer CoreWeave in briefing materials. No independent lab benchmark of the claim has been published as of this article.

What could limit how many Rubin GPUs ship in 2026?
Outside supply-chain estimates point to TSMC’s advanced packaging and N3 wafer capacity as the main bottleneck, with one analysis suggesting 2026 output could be capped in the range of 200,000 to 300,000 Rubin GPUs. Nvidia has not confirmed a specific production ceiling.