Nvidia’s next-generation AI rack just showed up in public benchmark data for the first time, and the numbers explain why the company keeps pulling its roadmap forward. On September 16, 2026, MLCommons published MLPerf Inference v6.1, and buried in a 486-result dataset was the debut of Vera Rubin NVL72, Nvidia’s successor to the Blackwell-based GB300 platform. In its first peer-reviewed showing, Vera Rubin posted up to 2.5x the DeepSeek-R1 throughput and up to 3.7x the Qwen3-VL throughput of Nvidia’s own current-generation GB300 NVL72 rack, according to figures Nvidia published alongside the MLCommons release.
The round did not belong to Nvidia alone. AMD used the same benchmark cycle to show off a 512-GPU Instinct MI355X cluster built with cloud provider Crusoe, and the broader submission list hit a participation record: 30 organizations, 120 systems, 486 individual results. For an industry spending hundreds of billions of dollars a year on AI data centers, MLPerf has become the closest thing to a neutral scoreboard, and this round gives buyers their first real look at what comes after Blackwell.
Vera Rubin’s First Public Benchmark Lands
Vera Rubin NVL72 entered MLPerf Inference v6.1 as a preview submission, meaning the hardware and software stack are not yet the shipping, generally available configuration Nvidia will eventually sell. MLCommons flags preview results separately from “available” results precisely because early silicon and early software drivers tend to improve before launch. Even with that caveat, a preview appearance from Nvidia’s flagship rack carries weight, because it is the first independently verified data point the market has had on Vera Rubin’s real-world inference performance.
Nvidia’s submission used a full 72-GPU NVL72 rack, with the individual Rubin accelerator listed in MLCommons’ records under the designation VR200. The DeepSeek-R1 runs relied on Nvidia’s TensorRT-LLM inference stack, while the Qwen3-VL-235B-A22B runs combined vLLM with Nvidia’s Dynamo inference framework. Nvidia also disclosed a larger four-rack configuration, 288 GPUs total, running DeepSeek-R1 offline at 99% scaling efficiency, a detail aimed squarely at hyperscale buyers who care less about single-rack peak numbers and more about how performance holds up once you chain racks together.
Cloud provider Nebius also submitted its own Vera Rubin NVL72 results, reporting 591,368 DeepSeek-R1 tokens per second in server mode and more than 1.1 million tokens per second on the gpt-oss 120B model in both server and offline scenarios. Those figures came from a separate system configuration and should not be read as identical to Nvidia’s own submission, but they confirm that more than one organization now has working Vera Rubin silicon in a lab, months ahead of general availability.
The Numbers: Vera Rubin NVL72 vs. GB300 NVL72
Nvidia’s own comparison, published alongside the MLCommons results, sets Vera Rubin against GB300 NVL72, its current top-of-line rack. It is not a comparison against GB200, B200, or H200, and no MLPerf v6.1 result table currently ties Vera Rubin directly to those older platforms. Treat any claim about Rubin’s advantage over H200 or B200 as unverified until Nvidia or MLCommons publishes that data directly.
| Workload | Scenario | Vera Rubin NVL72 | GB300 NVL72 | Rubin advantage |
|---|---|---|---|---|
| DeepSeek-R1 | Interactive | 652,750 tokens/s | 253,506 tokens/s | 2.58x |
| DeepSeek-R1 | Offline | 1,183,327 tokens/s | 679,740 tokens/s | 1.74x |
| DeepSeek-R1 | Server | 1,175,890 tokens/s | 596,944 tokens/s | 1.97x |
| Qwen3-VL-235B-A22B | Interactive | 1,306.6 queries/s | 349.3 queries/s | 3.74x |
| Qwen3-VL-235B-A22B | Offline | 2,392.7 samples/s | 1,305.0 samples/s | 1.83x |
| Qwen3-VL-235B-A22B | Server | 2,323.3 queries/s | 1,210.5 queries/s | 1.92x |
Two things jump out. First, the gap is widest on interactive scenarios, the workload class that most closely mirrors a live chatbot or agent waiting on a response, which suggests Vera Rubin’s architecture pays off most when latency budgets are tight. Second, the offline gap (where a system can batch requests without a latency ceiling) is the narrowest of the six rows, closer to 1.7x-1.8x than the headline 3.7x figure Nvidia leads with in its marketing. Read the “up to” language in Nvidia’s own materials literally: the biggest number is real, but it’s the ceiling, not the average.
Inside the Test: DeepSeek-R1 and Qwen3-VL Workloads
MLPerf Inference v6.1 leaned on two demanding, publicly available models to stress the new hardware. DeepSeek-R1, the reasoning-focused model that rattled the AI market when it first launched, measures raw token-generation speed, a proxy for how many chatbot or agent sessions a rack can serve at once. Qwen3-VL-235B-A22B, a large multimodal model, adds image and video understanding into the mix, measured in queries or samples per second rather than tokens.
Why Reasoning Models Changed the Benchmark
Reasoning models like DeepSeek-R1 generate far more intermediate tokens per answer than earlier chat models, since they work through steps before producing a final response. That makes raw tokens-per-second numbers less comparable across MLPerf rounds than they used to be, because a benchmark built around a reasoning model naturally rewards hardware with more memory bandwidth and faster interconnects, exactly the areas where Nvidia says Vera Rubin was redesigned.
Multimodal Load Adds a Second Dimension
Qwen3-VL’s inclusion also matters because it’s the first MLPerf Inference round to weight a multimodal workload this heavily alongside a pure-text reasoning model. Enterprise buyers building customer-support agents or document-processing pipelines increasingly need both capabilities on the same cluster, so a rack’s ability to hold up across both test types is arguably more useful to a buyer than either score in isolation.
AMD’s Counterpunch: 512 GPUs and Software Gains
AMD used the same round to argue that its existing hardware still has headroom. The company called v6.1 its broadest MLPerf submission yet, expanding from three model families in the prior v6.0 round to six across the MI355X, MI350X, and a new MI350P PCIe card, according to AMD’s own blog post on the results.
The MI350P appeared in a joint submission with Dell, using eight MI350P accelerators inside a PowerEdge XE7745 server, which Dell described as the only v6.1 submission to feature that specific card. But the bigger headline was software, not new silicon. On the identical eight-GPU MI355X configuration AMD used in the prior v6.0 round, the company reported a 28% jump in offline throughput and a 38% jump in server-mode throughput on the gpt-oss 120B model, plus a 70% gain on a single-stream Wan 2.2 video-generation test, all from ROCm software updates rather than a hardware refresh.
Then there was scale. Cloud provider Crusoe submitted results running gpt-oss 120B and DeepSeek-R1 across 512 MI355X GPUs, the largest single accelerator count ever submitted to an MLPerf Inference round. Crusoe’s numbers: roughly 5.75 million gpt-oss 120B tokens per second offline, 5.39 million tokens per second in server mode, and 2.40 million DeepSeek-R1 tokens per second in server mode. It’s a scale-out story rather than a per-chip efficiency story, and it lands at a moment when AMD is trying to convince hyperscalers that instinct GPUs cluster just as well as Nvidia’s NVLink-connected racks.
| Submission | Hardware | Workload | Result | Notable detail |
|---|---|---|---|---|
| Nvidia | Vera Rubin NVL72 (72 GPUs) | DeepSeek-R1, offline | 1,183,327 tokens/s | Preview category submission |
| Nvidia | Vera Rubin (4x NVL72, 288 GPUs) | DeepSeek-R1, offline | 99% scaling efficiency | Multi-rack scale-out test |
| Dell + AMD | MI350P (8 GPUs, PowerEdge XE7745) | GPT-OSS-120B, offline | ~48,000 tokens/s | Only v6.1 submission with MI350P |
| AMD | MI355X (8 GPUs, same as v6.0) | GPT-OSS-120B, offline/server | +28% / +38% vs. v6.0 | Software-only gain via ROCm |
| Crusoe | MI355X (512 GPUs) | GPT-OSS-120B, offline | ~5.75 million tokens/s | Largest accelerator count in MLPerf Inference history |
| Nebius | Vera Rubin NVL72 (preview) | gpt-oss 120B, server/offline | >1.1 million tokens/s | Independent lab confirmation of Rubin silicon |
A Record Field: 30 Submitters, 120 Systems, 486 Results
Beyond the Nvidia-versus-AMD storyline, MLCommons framed the whole round as a scale record. Thirty organizations submitted 120 distinct systems, generating 486 individual datacenter and edge results, according to MLCommons’ official release. The submitter list ranges well beyond chipmakers, spanning cloud infrastructure providers (CoreWeave, Oracle, Microsoft Azure, Lambda, Nebius), server builders (Dell, Supermicro, HPE, Fujitsu, Quanta), and even a solo academic contributor, Naeem Khoshnevis.
MLCommons did not publish a directly comparable organization or system count for the prior v6.0 round in its v6.1 materials, so a precise percentage increase can’t be verified here. What is clear is that this is one of the largest fields MLPerf Inference has run, and it’s the first round to add end-to-end retrieval-augmented generation (RAG) tests and edge-agentic benchmarks, categories built to mirror how companies are actually deploying AI today rather than isolated model tests.
Where Intel, Google, and the Rest Landed
Intel appeared in the round with Arc Pro B70 results, part of what MLCommons and outlet coverage flagged as one of several first-time peer-reviewed entries this cycle, alongside Vera Rubin, MI350P, and AMD’s Ryzen AI Max+ 395. Google was listed among the 30 submitting organizations, but the publicly available v6.1 result set does not identify a specific TPU v7 “Ironwood” score. That’s worth flagging clearly: Google’s name on the submitter list does not confirm Ironwood ran in this round, and no verified result should be attributed to it based on current public data.
That gap matters for anyone tracking the AI chip race broadly. Google has been steadily building out TPU capacity for its own Gemini workloads and cloud customers, and Nvidia has flagged both AMD and Google’s custom silicon efforts as the two rivals it takes most seriously in the accelerator market, a dynamic reflected in Nvidia’s ongoing push to lock in a large share of available high-bandwidth memory supply ahead of Vera Rubin’s launch.
Why the “Preview” Label Matters
MLCommons splits every round into “available” and “preview” categories for a reason. Available results come from hardware and software combinations any customer can buy and deploy today. Preview results, where Vera Rubin’s numbers sit, come from pre-launch silicon running on software that is still being tuned. Historically, Nvidia’s preview-to-available scores have moved in both directions between rounds: sometimes performance climbs as drivers mature, sometimes early preview numbers represent close to the ceiling.
The practical takeaway for IT buyers: treat the 2.5x and 3.7x figures as directional evidence of Vera Rubin’s architecture, not a guaranteed number you’ll see on day-one shipping hardware. Nvidia has not published a commercial launch date, transistor count, memory capacity, or per-unit TDP for Vera Rubin NVL72 alongside this MLPerf submission, so those specs remain unconfirmed even as the performance preview draws headlines.
How We Got Here: MLPerf’s Rise as the AI Hardware Scoreboard
MLPerf started in 2018 as a training benchmark built by academics and engineers frustrated with vendor-supplied numbers that couldn’t be compared apples-to-apples. Inference got its own dedicated benchmark track soon after, aimed at the far larger real-world question of how fast a chip serves an already-trained model to actual users. As generative AI shifted the industry’s spending from training runs to inference at scale, MLPerf Inference became the round that actually moves stock prices and procurement decisions, not the training benchmark.
The last two years pushed the benchmark toward rack-scale systems rather than single chips, mirroring how Nvidia, AMD, and cloud providers actually sell AI compute now. NVL72-class racks, tightly networked pools of dozens of GPUs acting as one giant accelerator, are the unit hyperscalers buy in bulk, and this round’s addition of RAG and agentic tests pushes MLPerf even further from lab benchmarks and closer to production workloads like customer support bots, coding assistants, and document search.
Nvidia vs. AMD: The Data Center AI Race in 2026
This MLPerf round captures a rivalry that has been building all year. AMD climbed past a trillion-dollar market cap in 2026 largely on the strength of its AI accelerator roadmap, and the company has been explicit that its pitch to hyperscalers rests as much on total cost of ownership and open ROCm software as on peak throughput. Nvidia, by contrast, is leaning on rack-scale integration, its NVLink interconnect, and now Vera Rubin’s early numbers to argue that its lead over merchant-silicon rivals isn’t closing.
Both arguments have merit depending on which number you weight. Nvidia’s per-rack advantage over its own prior generation looks real and substantial, particularly on latency-sensitive interactive workloads. AMD’s software-only 28-38% gains on unchanged hardware show it’s not standing still between chip generations, and Crusoe’s 512-GPU run proves Instinct clusters can scale to sizes that matter to hyperscale buyers. Neither company’s numbers, taken alone, settle which platform is the better buy. That decision increasingly comes down to price, supply, and how a given customer’s specific workload mix (chat-style reasoning versus multimodal versus RAG) maps onto each vendor’s strengths, an equation companies are already running as they weigh deals like India’s Yotta committing roughly $12 billion to 80,000 Nvidia Rubin GPUs.
What This Means for AI Infrastructure Spending
MLPerf results rarely move markets by themselves, but they do shape the procurement conversations happening inside every hyperscaler and enterprise IT department planning 2027 AI budgets right now. A credible 2-3x generational leap gives Nvidia ammunition to justify Vera Rubin pricing before the chip even ships, the same playbook it ran with Blackwell and Hopper before it. It also adds pressure on the memory supply chain: rack-scale AI systems eat high-bandwidth memory faster than fabs can currently produce it, a shortage already reshaping deals across the industry, including the data-center capex fight between Arm’s new AGI-focused CPU designs and AMD’s EPYC lineup.
For AMD, the message to investors and customers is different: don’t wait for new silicon to get better performance, because ROCm updates alone delivered close to 30-40% gains on hardware that’s already in data centers. That’s a cheaper pitch to make to a CFO than a full hardware refresh, and it’s consistent with AMD’s broader strategy of closing the gap with software velocity while its next-generation MI400-class chips finish development, a roadmap that overlaps with the broader wave of AMD and Nvidia next-gen GPU timelines now sliding toward 2027 and 2028.
The Fine Print: Software Stacks Skew the Comparison
Every number in this article comes with an asterisk worth repeating. Nvidia’s DeepSeek-R1 runs used TensorRT-LLM while its Qwen3-VL runs used vLLM plus Dynamo, two different software stacks tuned differently for different models. AMD’s gains came entirely from a ROCm software update on identical hardware. Precision settings, latency targets, power limits, and serving software all move MLPerf scores independently of the silicon underneath, which is exactly why MLCommons requires vendors to disclose their full software configuration alongside every submitted number.
That’s also why comparing Vera Rubin only to GB300, and not to GB200, B200, or H200, leaves a real gap in the public record. Until Nvidia or an independent submitter publishes that data, any claim about Rubin’s advantage over older Nvidia generations, or over Google’s TPU line, should be treated as speculation rather than benchmark fact. The most detailed independent write-up of the round makes a similar point: this cycle’s headline is software-driven optimization as much as it is new silicon.
What Industry Voices Are Saying
MLPerf Inference working-group chairs Miro Hodak and Frank Han published an analysis on September 17, 2026 framing the round around where the industry is actually putting its money: the record 30 submitters, 120 systems, and the arrival of agentic and end-to-end RAG benchmarks as a signal that inference testing is catching up to how companies deploy AI in production. Coverage from TechStrong AI on the same day characterized AMD’s contribution as software-driven improvement on existing MI355X hardware rather than a wholly new accelerator generation, pointing to the 28% and 38% GPT-OSS-120B gains and the 512-GPU Crusoe run as the round’s most concrete evidence. Neither MLCommons nor the vendors involved published a dedicated financial-analyst reaction alongside the technical results, so this round’s market read is still forming among traders and enterprise buyers parsing the numbers themselves.
Competitive Landscape at a Glance
| Vendor | Flagship in this round | Status | Standout number |
|---|---|---|---|
| Nvidia | Vera Rubin NVL72 | Preview (pre-launch) | Up to 3.7x GB300 on Qwen3-VL interactive |
| AMD | Instinct MI355X / MI350P | Available (shipping) | +38% server throughput via software alone |
| AMD (via Crusoe) | 512x MI355X cluster | Available (shipping) | Largest GPU count ever submitted to MLPerf Inference |
| Intel | Arc Pro B70 | First peer-reviewed entry | New to this benchmark class |
| Not specified in public results | Listed as submitter, no confirmed TPU score | Unconfirmed |
What Comes Next: Five Predictions
Nvidia will publish Vera Rubin NVL72 as an “available” (not preview) MLPerf submission within the next one to two rounds, most likely v7.0 or v7.1, once shipping hardware and finalized drivers are ready for outside testing. AMD will keep pushing software-driven gains on MI355X and MI350X through further ROCm releases rather than rushing a new chip to market before MI400-class silicon is ready. Expect a direct Vera Rubin-versus-GB200/B200/H200 comparison table to surface either from Nvidia’s own marketing or from a third-party lab within the next two to three months, closing the gap this round left open. Google will most likely submit a confirmed TPU v7 Ironwood result in a future round now that the company has re-entered the submitter list, especially as pressure builds to show its custom silicon keeps pace with merchant GPUs. And HBM supply, already tight enough to reshape deals across the industry, will keep shaping how fast either Vera Rubin or next-generation Instinct chips can actually reach customers regardless of what the benchmarks show.
The Buyer’s Bottom Line
For enterprise IT teams and cloud architects planning 2027 AI infrastructure, this round offers three concrete takeaways. First, Vera Rubin’s early numbers are real but preliminary, so don’t lock procurement decisions around preview-category scores that could shift before general availability. Second, AMD’s software-only gains prove that squeezing more performance out of already-purchased hardware remains a live option, which matters for budget cycles that can’t wait for next-generation chips. Third, workload mix now matters more than any single headline number: a cluster built for reasoning-heavy chat traffic and one built for multimodal RAG pipelines may not rank the same way on the vendor that looks best in a press release.
Frequently Asked Questions
What is MLPerf Inference v6.1?
It’s the latest round of MLCommons’ industry-standard AI inference benchmark, published September 16, 2026, covering how fast different chips and systems serve already-trained AI models to users. This round drew a record 30 submitting organizations and 486 individual results.
What is Nvidia Vera Rubin NVL72?
It’s Nvidia’s next-generation AI rack platform, the successor to the current Blackwell-based GB300 NVL72. It made its first public MLPerf appearance in this round as a preview submission, meaning the hardware and software are pre-launch and not yet generally available to customers.
How much faster is Vera Rubin than GB300?
Nvidia’s own published comparison shows up to 2.5x higher DeepSeek-R1 throughput and up to 3.7x higher Qwen3-VL throughput than GB300 NVL72, though the gains vary by scenario, ranging from roughly 1.7x on offline batch processing up to the 3.7x ceiling on interactive multimodal queries.
Is Vera Rubin faster than the H200 or B200?
That comparison hasn’t been published. Nvidia’s MLPerf v6.1 submission compares Vera Rubin only against GB300, so any claim about its advantage over H200, B200, or GB200 is currently unverified.
What did AMD show in this round?
AMD submitted results across six model families using MI355X, MI350X, and a new MI350P PCIe card, its broadest MLPerf submission to date. On identical MI355X hardware from the prior round, ROCm software updates alone delivered 28% higher offline and 38% higher server-mode throughput on the gpt-oss 120B model. Cloud provider Crusoe also ran a 512-GPU MI355X cluster, the largest accelerator count ever submitted to MLPerf Inference.
Did Google submit TPU results this round?
Google appears on MLCommons’ list of 30 submitting organizations, but no specific TPU v7 Ironwood result has been publicly identified in the v6.1 result set. Google’s presence on the submitter list should not be read as confirmation of a published Ironwood score.
When will Vera Rubin be available to buy?
Nvidia has not disclosed a commercial launch date, pricing, memory capacity, or TDP for Vera Rubin NVL72 alongside this MLPerf submission. Those details remain unconfirmed as of this benchmark round.
Why does MLPerf matter for AI infrastructure buying decisions?
MLPerf Inference is the closest thing the industry has to a standardized, audited benchmark for how AI chips perform on real workloads. With companies spending heavily on data-center buildouts, these results directly inform which chips hyperscalers and enterprises commit to before signing multi-year infrastructure deals, including projects like D-Matrix’s recent NVLink Fusion partnership with Nvidia.



