Nvidia just opened the door to its own server racks for a rival chip, and the first company through it is a Santa Clara startup most gamers have never heard of. On September 10, 2026, d-Matrix said it will adopt Nvidia’s NVLink Fusion interconnect to plug its upcoming Raptor inference chips directly into Nvidia’s rack-scale AI infrastructure, according to Reuters. The deal turns Nvidia’s own NVL144 MGX rack architecture into a host for someone else’s silicon, a move that would have sounded unthinkable a few years ago from a company that has spent a decade building a walled garden around its GPUs.
The timing lands in the middle of an inference capacity crunch that has defined this year’s hardware news cycle. AI labs are shipping models faster than data centers can serve them, and every hour of GPU downtime now carries a real dollar cost. d-Matrix’s pitch is narrow but pointed: let its memory-heavy Raptor chips handle the token-by-token decode phase of large language model inference while Nvidia’s own GPUs handle the heavier prefill phase, all inside the same rack, talking over the same interconnect fabric. Nvidia’s own newsroom confirmed the arrangement, describing d-Matrix as joining “a growing roster of ecosystem partners” building on NVLink Fusion, according to the NVIDIA Newsroom.
What d-Matrix and Nvidia Actually Announced
The core of the announcement is technical integration, not an acquisition or an equity deal. d-Matrix confirmed it has designed Raptor, its next-generation inference chip, “from the ground up for integration with NVIDIA NVLink Fusion,” according to the company’s own announcement page. As an NVLink Fusion partner, d-Matrix will incorporate its next-generation inference XPUs into Nvidia’s latest rack reference architecture, a design that also includes Nvidia Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking, per the same release.
Raptor is expected to tape out, meaning finalize its physical chip design, before the end of 2026. Systems that combine Raptor with Nvidia’s MGX rack are targeted for initial availability in Q4 2027, more than a year out. That gap matters: this is a product-roadmap commitment, not a shipping system, and buyers evaluating either company’s near-term hardware plans should treat 2027 as the operative date, not 2026.
d-Matrix CEO and cofounder Sid Sheth framed the deal around deployment speed rather than raw performance. “With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference,” Sheth said, according to Nvidia’s official blog. He added a line that captures the industry’s current bottleneck: “Demand for inference is soaring, but capital, time and energy remain finite,” also per Nvidia’s blog post on the partnership.
Inside Raptor: The Chip d-Matrix Is Betting On
Raptor is not d-Matrix’s first product. It follows Corsair, the company’s current-generation inference accelerator, which the company says delivers up to 30,000 tokens per second on a Llama 70B model with roughly 2 milliseconds of latency per token, according to a December 2025 analysis republished by Financial Content. Corsair’s whole design premise is in-memory computing: keeping model weights close to compute so inference doesn’t stall waiting on data movement.
Raptor pushes that same idea further with a 4-nanometer process and a 3D-stacked DRAM technology the company calls 3DIMC, validated through test silicon d-Matrix calls Pavehawk. At the Hot Chips conference, d-Matrix disclosed that each Raptor card will carry 32GB of ultra-fast 3D-stacked DRAM and deliver 100TB/s of memory bandwidth per card, a figure The Register calculated at roughly 4.5 times the memory bandwidth of Nvidia’s own Rubin GPU. By the end of 2027, d-Matrix expects to offer systems with up to 144 Raptor accelerators tied together on a single all-to-all NVLink fabric, per the same Register report.
Sheth described the practical upside of that rack compatibility in plain terms: “The beauty of this solution is we take the Raptor trays, plug them into the same NVL144 MGX rack architecture, which is widely deployed across many datacenters, and we get instant access to those datacenters,” he told The Next Platform. That is the real value proposition here: d-Matrix doesn’t need customers to build new data centers or new cooling loops, it just needs them to swap trays inside racks they already have.
What NVLink Fusion Actually Is
NVLink Fusion is Nvidia’s modular interconnect architecture, first unveiled at Computex 2025, that lets non-Nvidia accelerators, CPUs, and networking hardware talk to Nvidia GPUs over the same high-bandwidth NVLink fabric that normally only connects Nvidia’s own chips. Nvidia’s fifth-generation NVLink platform, used in GB200 NVL72 and GB300 NVL72 racks, delivers 1.8TB/s of bandwidth per GPU, which Nvidia describes as 14 times faster than PCIe Gen5, according to the company’s own May 2025 announcement. Reuters’ September 10 report on the d-Matrix deal notes Raptor will connect through Nvidia’s newer sixth-generation NVLink fabric and slot into the MGX reference rack design.
d-Matrix is not the first company to sign on. Launch partners named when NVLink Fusion debuted included MediaTek, Marvell, Alchip Technologies, Astera Labs, Synopsys, and Cadence, with Fujitsu and Qualcomm separately confirming plans to pair custom CPUs with Nvidia GPUs using the same scale-up technology, according to EfficientlyConnected. Nvidia has also backed the strategy with cash: it invested $2 billion into Marvell in April 2026 specifically to deepen the NVLink Fusion ecosystem, even though Marvell’s biggest customers are simultaneously trying to build chips that compete with Nvidia, as Tom’s Hardware reported.
That tension is the story underneath the story. Nvidia is opening its rack architecture to chips explicitly designed to reduce how much Nvidia silicon a customer needs to buy, and it’s doing so on purpose. The read from most analysts is that Nvidia would rather own the connective tissue of the AI data center than fight every custom-chip effort head-on. If a hyperscaler is going to buy some non-Nvidia silicon anyway, Nvidia wants that silicon still riding on NVLink, still buying Nvidia switches, DPUs, and networking gear around it.
Why This Matters More Than It Looks
Counterpoint Research has described NVLink Fusion directly as “Nvidia’s Response to UALink,” the open interconnect standard backed by AMD, Broadcom, and a coalition of other chipmakers that want a vendor-neutral alternative to Nvidia’s proprietary fabric, according to Counterpoint’s analysis. UALink’s pitch to chipmakers is independence: build once, connect to any vendor’s rack. NVLink Fusion’s pitch is different but arguably stronger in the short term, since it hands partners immediate access to Nvidia’s already-deployed installed base of racks, cooling systems, and networking, rather than asking data center operators to adopt a newer, less proven standard.
Nvidia has been busy on other fronts too, including its own reported $12.9 billion move to acquire Hugging Face, another sign the company is buying and partnering its way around every part of the AI stack it doesn’t already control outright. For d-Matrix specifically, the calculation looks straightforward. The company raised $275 million in a round that valued it at roughly $2 billion, reported by DataCenterDynamics in July 2026. That is real money, but it is nowhere near what it would cost to convince data center operators to build entirely new rack infrastructure around a chip most of them have never deployed. Riding into Nvidia’s already-installed MGX racks skips that cold-start problem entirely.
Yahoo Finance’s coverage of the deal points to a quieter beneficiary: Astera Labs, the connectivity chipmaker whose retimers and fabric switches sit inside NVLink Fusion deployments regardless of which accelerator vendor wins a given contract, according to the Yahoo Finance analysis. Every additional NVLink Fusion partner is, in effect, another customer for the physical connectivity layer Astera Labs supplies, independent of whether Nvidia, d-Matrix, or some other accelerator ends up doing the actual math.
The Split-Inference Pitch: Prefill vs. Decode
The technical logic behind pairing Raptor with Nvidia GPUs rests on a split in how large language model inference actually works. Every inference request runs through two distinct phases. Prefill processes the entire input prompt at once, a compute-heavy burst that benefits from raw GPU throughput. Decode then generates the response one token at a time, a phase that is far more sensitive to memory bandwidth and latency than to raw compute power.
Nvidia’s Vera Rubin platform is built to dominate prefill. d-Matrix is positioning Raptor, with its 100TB/s of per-card memory bandwidth, to take over decode instead, according to analysis from FourWeekMBA. Splitting the two phases across specialized hardware inside a single rack, connected by a fast enough fabric that the handoff doesn’t create its own bottleneck, is the bet both companies are making. If it works at scale, operators get lower cost per token on decode-heavy workloads like chatbots and coding assistants without giving up prefill throughput on the compute-heavy side.
Why Inference Economics Are Driving This Deal
Training gets the headlines, but inference is where AI companies actually burn money at scale, every single day, across millions of user requests. A chip that shaves milliseconds off decode latency or doubles tokens-per-dollar on that phase changes the unit economics of running a chatbot or coding assistant at scale. That is the gap d-Matrix is chasing, and it’s the same gap that has pulled in a wave of inference-focused chip startups over the past two years.
Nvidia’s Widening Circle of “Frenemies”
d-Matrix joining NVLink Fusion is the latest data point in a pattern that has been building through 2026: Nvidia’s biggest customers and its most direct rivals are increasingly the same companies, and Nvidia keeps finding ways to stay in the middle of the transaction anyway. The company just posted a $96.2 billion quarter built almost entirely on GPU demand, even as cloud providers and hyperscalers keep funding custom silicon meant to reduce their dependence on exactly that GPU demand, a dynamic shattered.io covered in detail after Nvidia’s latest earnings. Separately, Nvidia’s customers building rival chips has become its own storyline this year, one this site has also tracked directly.
The NVLink Fusion strategy is Nvidia’s answer to that contradiction: rather than losing those deals outright, it tries to make sure any custom silicon a hyperscaler builds still has to move data through Nvidia’s fabric, still needs Nvidia’s switches, and still gets deployed inside Nvidia’s rack reference designs. It’s a hedge that only a company with Nvidia’s market position could plausibly pull off, since it requires partners to want access to an installed base large enough that plugging into it beats building an independent one.
NVLink Fusion Partners: Who’s Building What
| Partner | Role in NVLink Fusion | Status as of Sept. 2026 |
|---|---|---|
| d-Matrix | Custom inference XPU (Raptor) | Announced Sept. 10, 2026; racks targeted Q4 2027 |
| MediaTek | Custom silicon / SoC design | Launch partner since Computex 2025 |
| Marvell | Custom ASIC and connectivity | Launch partner; received $2B Nvidia investment in April 2026 |
| Qualcomm | Custom CPU integration with Nvidia GPUs | Confirmed integration plans, 2025 |
| Fujitsu | Custom CPU integration with Nvidia GPUs | Confirmed integration plans, 2025 |
| Astera Labs | Connectivity chips (retimers, fabric switches) | Launch partner; flagged as a quiet beneficiary of new deals |
| Alchip Technologies | ASIC design services | Launch partner since Computex 2025 |
Raptor vs. Corsair vs. Nvidia Rubin: The Numbers
| Chip | Maker | Memory Bandwidth | Notable Spec | Availability |
|---|---|---|---|---|
| Corsair | d-Matrix | Not disclosed at card level | 30,000 tokens/sec on Llama 70B, ~2ms/token latency | Current generation |
| Raptor | d-Matrix | 100TB/s per card (32GB 3D-stacked DRAM) | 4nm process; up to 144 cards per NVLink fabric | Tape-out late 2026; racks Q4 2027 |
| Rubin GPU | Nvidia | ~22TB/s per card (per The Register’s comparison) | Reference point Raptor’s bandwidth is measured against | 2026-2027 rollout |
| NVLink (5th-gen platform) | Nvidia | 1.8TB/s per GPU fabric bandwidth | 14x PCIe Gen5 bandwidth; up to 72 GPUs per NVL72 rack | Shipping (GB200/GB300 NVL72) |
The bandwidth comparison in that table is the headline number The Register calculated: Raptor’s 100TB/s of per-card memory bandwidth works out to roughly 4.5 times Nvidia’s own Rubin GPU. That’s a memory bandwidth claim specifically, not a claim about overall inference throughput or total system performance, and it says nothing about cost per card or power draw, figures neither company has disclosed yet.
How This Fits the Broader Memory and GPU Market
Raptor’s memory-centric design lands at a moment when memory itself has become the scarcest input in AI hardware. High-bandwidth memory yields and allocation have become a bottleneck across the industry this year, a trend shattered.io has tracked as HBM4 supply ramps for Nvidia’s own Rubin platform. A chip like Raptor, built specifically to squeeze more bandwidth out of 3D-stacked DRAM rather than relying purely on HBM stacks, is partly a bet that alternative memory architectures can sidestep some of that scarcity.
It also lands as GPU rental pricing keeps shifting. Cloud H100 rental rates have fallen roughly in half over the past year even as newer B200 capacity commands a premium, a dynamic this site broke down in its coverage of current cloud GPU pricing. Specialized inference silicon like Raptor is explicitly trying to carve out a cost advantage in that gap, targeting workloads where a general-purpose GPU is doing more (and costing more) than the job strictly requires.
Historical Context: From Closed Fabric to Semi-Open Ecosystem
NVLink began life in 2016 as a way to connect Nvidia GPUs to each other and to IBM Power CPUs, a strictly closed, Nvidia-to-Nvidia (and select-partner) fabric. For most of the past decade, NVLink access was a locked door: it connected Nvidia chips to Nvidia chips, full stop. NVLink Fusion, announced at Computex 2025, cracked that door open for the first time, letting outside silicon into a fabric Nvidia had spent years hardening as a competitive moat.
The shift mirrors what happened with PCIe adoption decades earlier: an initially closed or fragmented connectivity standard eventually opens up once enough of the market demands interoperability. The difference here is that Nvidia isn’t ceding control by opening NVLink Fusion, it’s extending its reach. Every chip that adopts NVLink Fusion still depends on Nvidia’s switches, Nvidia’s rack designs, and increasingly, in Marvell’s case, Nvidia’s own investment capital.
Competitive Landscape: NVLink Fusion vs. UALink vs. Broadcom’s Approach
UALink remains the clearest structural alternative. Backed by AMD, Broadcom, and other members of a broader industry consortium, UALink is designed as an open standard that any chipmaker can build to without needing Nvidia’s cooperation or Nvidia’s rack designs. The appeal is independence. The tradeoff is that UALink-based systems don’t inherit an existing, massive installed base the way NVLink Fusion partners do, since Nvidia’s racks are already deployed by the thousands across hyperscale data centers worldwide.
Broadcom, meanwhile, has built its custom-XPU business largely around direct hyperscaler relationships, most notably supplying Google’s TPU lineage and other custom accelerator programs, rather than trying to plug into someone else’s fabric. That approach sidesteps dependency on Nvidia entirely but requires each hyperscaler to build and validate its own end-to-end rack architecture, a heavier lift than adopting NVLink Fusion. For a startup like d-Matrix, without Google-scale resources to build an independent rack ecosystem, riding Nvidia’s existing infrastructure is the more capital-efficient path, even if it means playing inside Nvidia’s fabric rather than outside it.
What Analysts and Investors Are Watching
Reuters framed the announcement as part of a broader pattern of intensifying demand for AI compute, noting that custom AI chips are increasingly being built to plug directly into Nvidia’s existing data-center systems rather than compete against them from outside. That framing lines up with how Nvidia itself talks about the strategy: expanding what counts as “Nvidia infrastructure” even when the chip actually doing the inference math isn’t Nvidia’s own silicon.
Investors parsing the deal have focused less on d-Matrix directly, since it remains privately held, and more on which public companies benefit from every additional NVLink Fusion partner. Astera Labs’ role supplying the physical connectivity layer inside these deployments is the clearest example, since its retimers and fabric switches are needed regardless of which accelerator vendor a given customer ultimately picks for a given rack.
5 Predictions for Where This Goes Next
- Expect more inference-focused startups to announce NVLink Fusion integrations over the next two quarters, following the same playbook: ride Nvidia’s installed base rather than build a competing one.
- d-Matrix’s Q4 2027 rack availability target will likely slip, as most first-generation rack-scale hardware integrations in this industry tend to run behind their initial public timelines.
- Astera Labs and other connectivity suppliers embedded in NVLink Fusion racks stand to see steady demand growth regardless of which accelerator vendor wins any individual inference contract.
- UALink’s backers will use deals like this one to argue harder for an open standard, framing NVLink Fusion as deeper vendor lock-in dressed up as openness.
- Nvidia will keep making direct equity investments, following the Marvell pattern, in NVLink Fusion partners it considers strategically important enough to keep close, even when those same partners’ customers are trying to reduce Nvidia dependence.
The Bigger Picture for Data Center Buyers
For the hyperscalers and cloud providers actually buying this hardware, the d-Matrix deal is a signal that mixing accelerator vendors inside a single rack is becoming a supported, sanctioned pattern rather than an unofficial workaround. That matters operationally. It means data center architects can plan for heterogeneous racks, Nvidia GPUs for compute-heavy prefill, specialized chips like Raptor for latency-sensitive decode, without betting an entire infrastructure refresh on a single vendor’s roadmap holding up.
It also raises the bar for what a would-be Nvidia competitor needs to offer. A chip that can’t plug into NVLink Fusion, and therefore can’t ride Nvidia’s already-deployed rack infrastructure, now has to win on cost or performance alone, without the advantage of an existing installed base. That’s a steeper climb than the one d-Matrix just signed up for.
Frequently Asked Questions
What is NVLink Fusion?
NVLink Fusion is Nvidia’s modular interconnect architecture, first announced at Computex 2025, that allows non-Nvidia chips, including custom CPUs, ASICs, and inference accelerators, to connect to Nvidia GPUs and rack infrastructure over the same high-bandwidth NVLink fabric Nvidia normally reserves for its own hardware.
What is d-Matrix’s Raptor chip?
Raptor is d-Matrix’s next-generation AI inference chip, built on a 4-nanometer process with 3D-stacked DRAM memory technology. Each card is expected to deliver 32GB of memory and 100TB/s of memory bandwidth, according to disclosures at the Hot Chips conference reported by The Register.
When will d-Matrix Raptor systems be available?
Raptor is expected to complete its final chip design, or tape out, before the end of 2026. Rack-scale systems that integrate Raptor into Nvidia’s MGX reference architecture are targeted for initial availability in Q4 2027.
Is Nvidia investing money in d-Matrix?
No direct equity investment from Nvidia into d-Matrix has been reported as part of this announcement. This is described in all available reporting as a technical integration partnership. Nvidia has made direct investments in other NVLink Fusion partners, including a $2 billion investment in Marvell reported in April 2026.
How much funding has d-Matrix raised?
d-Matrix raised $275 million in a round that valued the company at roughly $2 billion, according to DataCenterDynamics’ report from July 2026.
What is the difference between NVLink Fusion and UALink?
NVLink Fusion is Nvidia’s proprietary interconnect, opened selectively to partner chipmakers. UALink is an open industry standard backed by AMD, Broadcom, and other companies that lets any member build compatible hardware without depending on Nvidia’s rack designs or cooperation. Counterpoint Research has described NVLink Fusion as effectively Nvidia’s competitive response to UALink’s emergence.
Why does memory bandwidth matter so much for AI inference chips?
The decode phase of large language model inference, where a model generates a response one token at a time, is limited more by how fast data can move between memory and compute than by raw processing power. Chips with higher memory bandwidth, like Raptor’s 100TB/s per card, can generate tokens faster and at lower latency during this phase.
Which other companies have adopted NVLink Fusion?
Launch partners named when NVLink Fusion debuted in 2025 included MediaTek, Marvell, Alchip Technologies, Astera Labs, Synopsys, and Cadence, with Fujitsu and Qualcomm separately confirming plans to integrate custom CPUs with Nvidia GPUs using the same technology. d-Matrix is the most recent company to join, announced September 10, 2026.




