Nvidia's Grip on AI Data Centers Faces a Challenge From an Open Memory Standard
Two of the world's largest memory chip makers are betting that an open-standard technology called CXL (Compute Express Link) will become essential to the next generation of AI data centers, even as Nvidia's dominance in the field creates significant headwinds for its adoption. At the Future of Memory and Storage conference in Santa Clara this week, Samsung Electronics and SK Hynix unveiled major advances in CXL-based systems designed to solve a critical bottleneck in large language model (LLM) inference, where AI systems must process and remember long conversations or documents.
What Problem Is CXL Trying to Solve in AI?
As AI moves beyond simple learning tasks into the era of LLM inference, where models must understand long contexts, demand for KV cache (Key Value Cache) is surging. LLM-based AI agents temporarily store prior conversation history in KV cache to speed up inference, and as conversations grow longer and user counts rise, KV cache memory requirements can balloon to hundreds of gigabytes.
Samsung Electronics demonstrated the problem and its proposed solution at the conference. In a test with a conventional 512GB DRAM configuration, KV cache demand exceeded available memory capacity and performance degraded significantly. However, when Samsung added a 1TB CXL memory pool to the system, it sustained KV cache demand and kept performance stable. SK Hynix, working with Marvell, took a different approach by embedding computation directly into memory. Its CMM-Ax module achieved throughput up to 5.5 times higher than a single GPU and 3.6 times higher than a dual-GPU setup in ultra-long-context LLM environments.
Why Isn't CXL Already Everywhere if It Works?
The answer involves a key industry player with little interest in seeing CXL spread: Nvidia. Nvidia connects GPUs and CPUs through its proprietary NVLink standard, now in its sixth generation, which supports 3.6 terabytes per second of bandwidth per GPU. This enables tens to hundreds of GPUs in Blackwell and Rubin-based data centers to work together like a high-performance supercomputer.
Nvidia's strength in the AI data center market lies in locking customers into its ecosystem on the back of dominant GPU performance. The company has built an exclusive programming ecosystem around CUDA, its proprietary parallel computing platform. The spread of CXL, an open interconnect standard, would inevitably affect Nvidia's position. While Nvidia joined the CXL consortium in 2019, largely to monitor future customer demand and keep a check on the ecosystem, its enthusiasm for the standard remains limited.
Intel launched the CXL consortium in 2019 to expand the ecosystem around its Xeon data center CPUs, drawing in global cloud hyperscalers and memory companies. However, Nvidia's dominance in AI accelerators has made it difficult for open standards to gain traction in the data center market.
Are Memory Makers Pushing Back Against Nvidia's Control?
Samsung and SK Hynix are making their case directly. At the conference, Samsung engineers So Jin-in and Lee Ho-gyun presented findings titled "The Revival of CXL in the AI Era," directly referencing NVLink as a comparison point.
"A few years ago, some industry reports claimed that CXL was finished in the AI era, citing bandwidth limitations compared with interconnects like NVLink. CXL is now emerging as a core solution by leveraging its low-latency characteristics alongside large-capacity memory pooling and sharing technology," the two said.
So Jin-in, Senior Director of Memory System Architecture at Samsung Electronics, and Lee Ho-gyun, Director of DRAM Solution Engineering at Samsung Electronics
The argument reflects a broader industry shift. As AI inference workloads become more complex and memory-intensive, the idea of a shared memory reservoir measured in petabytes becomes increasingly attractive. Both Samsung and SK Hynix have been developing CXL technology for years. Samsung became the first in the industry to develop CXL-based DRAM technology in 2021, and SK Hynix followed the very next year with its own CXL memory samples.
What Other Memory Technologies Are Emerging?
Beyond CXL, both companies showcased another next-generation technology at the conference: PIM (Processing-In-Memory), which enables memory chips to perform computation in addition to storage. This represents a fundamental shift in computer architecture, where the processor handling computation and the memory storing data have traditionally been physically separate.
By designing computational circuits inside the storage cells themselves, memory can handle basic operations on its own. This reduces the amount of data moving between components, boosting speed and cutting power consumption. Samsung Electronics brought its LPDDR5X-PIM (low-power DRAM with PIM) to the conference and won the FMS Best of Show Award. SK Hynix also outlined its development roadmap for LPDDR6-PIM.
How Are Memory Makers Positioning These Technologies?
- Edge AI Applications: PIM technology targets the edge AI market, enabling devices such as on-device AI to operate without routing through a central cloud server, allowing high-performance AI to run at low power even without large servers.
- Long-Context LLM Support: CXL memory pooling addresses the specific challenge of handling massive KV cache demands in ultra-long-context language models, where traditional DRAM alone cannot meet capacity requirements.
- Hybrid Architectures: SK Hynix's CMM-Ax combines CXL connectivity with processing capabilities, embedding 16 CPUs in Marvell's CXL controller to amplify the synergy between memory and computation.
The conference also revealed developments in NAND flash technology. SK Hynix, which leads the HBM (High Bandwidth Memory) market by share, joined with SanDisk to publish the first formal HBF (High Bandwidth Flash) standard specification, covering capacity, bandwidth and connectivity. The idea is that HBF needs to sit within the existing memory hierarchy between GPU HBM, DRAM and SSDs to provide additional capacity at a relatively lower cost than DRAM.
The broader story reflects a tension in the AI infrastructure market. While Nvidia's CUDA ecosystem and NVLink standard have created an unmatched competitive moat, the memory demands of modern AI are forcing the industry to explore alternatives. Whether CXL and other open standards can gain meaningful adoption against Nvidia's entrenched position remains an open question, but the momentum from the world's two largest memory chip makers suggests the conversation is far from over.