NVIDIA's Inference Chip Strategy Just Shifted: Why It's Betting on d-Matrix Over Going Solo
NVIDIA announced a multi-year partnership with d-Matrix that brings specialized inference chips directly into its core AI infrastructure, marking a strategic pivot in how the company manages competitors in the booming inference market. Rather than building everything in-house, NVIDIA is now integrating d-Matrix's Raptor inference processors into its latest MGX rack architecture through NVLink Fusion, a connectivity standard that lets different chip makers work seamlessly together.
What Makes This Partnership Different From NVIDIA's Usual Playbook?
For years, NVIDIA dominated AI by controlling both the chips and the ecosystem around them. This deal represents something new: NVIDIA is essentially inviting a specialized competitor inside the tent. d-Matrix will integrate its Raptor inference XPUs (specialized processors designed for AI inference, the process of running trained models to generate outputs) into NVIDIA's MGX racks alongside NVIDIA's own Vera CPUs, BlueField-4 data processing units, and ConnectX-9 networking components.
The partnership reflects a pragmatic reality in the inference market. NVIDIA already owns a decode specialist through its Groq licensing deal, which produces the Groq 3 LPX rack capable of delivering 11,000 tokens per second on smaller models. But that system uses SRAM memory, which limits it to compact models around 2 trillion parameters. d-Matrix's Raptor uses 3D-DRAM stacked memory instead, giving it 2.3 terabytes of capacity compared to Groq's 128 gigabytes, making it suitable for much larger models.
How Does the Technical Architecture Actually Work?
The Raptor rack design stacks 144 XPU processors across 18 compute trays, delivering roughly 1,000 tokens per second per user on 3-trillion-parameter models with 1 million token context windows (the amount of previous conversation the model can remember). This represents a fundamentally different approach than NVIDIA's Groq partnership, which prioritizes speed on smaller models over capacity on larger ones.
The heterogeneous disaggregation model splits workloads intelligently: NVIDIA GPUs handle the compute-intensive prefill phase (preparing the model to generate), while d-Matrix XPUs accelerate the latency-sensitive decode phase (actually generating tokens). This division of labor has already been validated in real deployments. Gimlet Labs demonstrated a 2x to 5x end-to-end speedup on interactive workloads using this split, while Parasail achieved up to 10x faster interactive inference with 3x better energy efficiency.
Why Should Companies Care About This Inference Chip Shift?
The inference market is exploding. Futurum's 2026 forecast projects agent and reasoning-focused inference silicon will grow from $35.9 billion in 2025 to $546 billion by 2030, surpassing pre-training as the largest AI workload segment by 2027. The XPU sub-market that d-Matrix competes in is expected to expand from $37.4 billion to $237.2 billion over the same period.
For hyperscalers and AI labs, this partnership means they can now deploy frontier-scale models with massive context windows without being locked into a single vendor's approach. The MGX supply chain is already mature and widely deployed, giving d-Matrix immediate access to qualified procurement channels that competitors like Cerebras and AMD lack.
Steps to Understanding the Inference Chip Landscape
- Prefill vs. Decode: Prefill is the compute-heavy phase where models process input and prepare to generate. Decode is the latency-sensitive phase where models generate one token at a time. Different hardware excels at each task, which is why heterogeneous systems are becoming standard.
- Memory Architecture Matters: SRAM-based systems like Groq 3 LPX are extremely fast but capacity-limited. DRAM-based systems like Raptor sacrifice some speed for dramatically larger capacity, enabling larger models and longer context windows.
- Ecosystem Integration: Being part of NVIDIA's MGX ecosystem means access to proven supply chains, software compatibility, and customer relationships. Standalone chips, no matter how fast, struggle without this infrastructure.
When Will This Actually Ship, and What Are the Risks?
The timeline reveals the distance between announcement and revenue. Raptor has not yet taped out (the final step before manufacturing), with tape-out expected before the end of 2026. Initial availability of Raptor XPUs in MGX racks is projected for Q4 2027, roughly two years from now. Meanwhile, the Groq 3 LPX entered full production this year, giving it a significant head start.
d-Matrix has traded a measure of independence for entry into the ecosystem it once positioned itself against. The company's Corsair platform, currently in production, has proven the decode acceleration concept through deployments at Gimlet Labs and Parasail. But Raptor's performance projections remain pre-silicon simulations from a chip that has not yet been manufactured.
"Being integrated into NVIDIA's latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform," said Sid Sheth, Founder and CEO of d-Matrix.
Sid Sheth, Founder and CEO of d-Matrix
What Does This Mean for Other Inference Chip Makers?
The partnership illustrates NVIDIA's strategy for managing inference challengers: invite them into the rack and monetize everything around them. Cerebras projects up to 5,000 tokens per second on future wafer-scale systems, and AMD acquired Taalas for hardcoded inference silicon, but neither company has plugged into the MGX supply chain that hyperscaler procurement already qualifies.
The technical validation is compelling. d-Matrix's ISCA 2026 paper on early Raptor silicon showed roughly 100 terabytes per second of memory bandwidth per card at 0.45 picojoules per bit of input-output energy, about 6 times more efficient than HBM3 memory. On speech models like Whisper and Canary, Raptor delivered 4.71 times higher throughput per card than HBM-based designs and 2.44 times higher than SRAM-based designs.
This partnership signals that the future of AI inference is not about one company controlling everything. Instead, it is about specialized chips working together within a trusted ecosystem. NVIDIA maintains control of the platform and the supply chain, while d-Matrix gets access to hyperscaler customers. The real winners are the companies deploying these systems, who now have options for serving both small, fast models and massive, long-context models from a single integrated rack.