Logo
FrontierNews.ai

Former Groq Engineers Launch $875M Startup to Rethink How AI Servers Store and Access Data

A startup led by former Groq and Lambda executives just secured $875 million to challenge how the AI industry thinks about inference hardware, betting that mobile-grade memory can outperform the specialized chips currently dominating data centers. Positron AI, founded in April 2023 by neocloud veterans Mitesh Agrawal, Thomas Sohmers, and Edward Kmett, is designing what it calls "memory-first inference systems" that prioritize data flow over raw computing speed.

The funding came in two tranches: a $375 million Series C round co-led by Atreides Management and Valor Equity Partners, followed by a $500 million Series C-1 round led by New Enterprise Associates and Netscape co-founder Jim Clark. The startup is now valued at around $5 billion, with backing from Qatar's sovereign wealth fund, Cisco Investments, and other major investors.

Why Does Memory Bandwidth Matter More Than You Might Think?

The core insight behind Positron's approach is deceptively simple: modern AI models are bottlenecked not by how fast chips can calculate, but by how quickly they can move data in and out of memory. Think of it like a highway where the calculation lanes are wide open, but the on-ramps and off-ramps are clogged. Positron's custom silicon, called Asimov, claims to achieve more than 90 percent of available memory bandwidth, compared to just under 30 percent for graphics processing units (GPUs) running the same models.

Instead of using high-bandwidth memory (HBM), a specialized and expensive type of memory designed for data centers, Positron opted for LPDDR5X (low-power double data rate 5X), a memory technology originally designed for smartphones and edge devices. This unconventional choice sidesteps the current supply constraints and power limitations plaguing HBM while delivering the massive data throughput that large language models (LLMs) require.

How Does Positron's Hardware Stack Actually Work?

  • Memory Capacity: Each Asimov chip can hold between 277 gigabytes and 2.3 terabytes of data, allowing servers to store enormous models without constantly fetching data from external storage.
  • Scalability: Multiple Asimov chips can be interconnected into clusters containing up to 16,384 chips, with chip-to-chip bandwidth reaching around 128 terabits per second to keep data flowing smoothly across the system.
  • Server Configuration: The Titan inference server platform houses either four or eight Asimov chips and can support up to 32 trillion parameters per server, enough to run the largest frontier models currently in development.
  • Cooling Options: Titan servers are designed for both air-cooled and liquid-cooled deployments, giving data centers flexibility in how they manage heat and power consumption.

The founding team brings deep expertise from the cloud infrastructure world. Agrawal spent more than seven years at Lambda, a cloud platform for AI workloads, leading AI and machine learning GPU revenue and operations. Kmett worked as a hardware architect at Lambda before joining Groq, the language-processing unit (LPU) developer that was largely acquired by Nvidia in an acqui-hire. Sohmers also came from Groq, where he served as head of technology and architecture.

"The Positron inference architecture balances compute, the enormous required memory bandwidth, and extraordinarily large context and weight storage. Its optimized power consumption, cost, and density is near ideal for inference requirements of the next evolution of frontier large language models with trillions of parameters," said Jim Clark, Netscape co-founder and investor.

Jim Clark, Netscape Co-founder

What's the Timeline for Getting These Chips Into Production?

Positron is moving fast. The $875 million funding will support three critical milestones: tapeout of the Asimov chip, production ramp of Titan servers, and securing LPDDR5X memory capacity. Asimov is set for tapeout on Taiwan Semiconductor Manufacturing Company's (TSMC) advanced three-nanometer process at the end of 2026, with production ramping in the second half of 2027.

Mitesh Agrawal, Positron's CEO, emphasized the importance of speed in a market where inference hardware is becoming increasingly competitive. "Speed matters in this market, both in how quickly we ship new generations of silicon and in how quickly they reach customers," he stated. "Deploying Atlas at scale taught us an enormous amount about what inference customers actually need, and we have carried those lessons directly into Asimov and Titan. Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us".

Mitesh Agrawal, Positron's CEO

"Speed matters in this market, both in how quickly we ship new generations of silicon and in how quickly they reach customers. Deploying Atlas at scale taught us an enormous amount about what inference customers actually need, and we have carried those lessons directly into Asimov and Titan," explained Mitesh Agrawal, CEO of Positron AI.

Mitesh Agrawal, CEO at Positron AI

Why Does This Challenge Nvidia's Dominance?

Positron's approach represents a fundamental rethinking of inference hardware architecture. While Nvidia's GPUs excel at training models (the computationally intensive phase), inference (running trained models to generate outputs) has different requirements. Inference workloads care less about peak compute performance and more about moving massive amounts of data efficiently. By using mobile-grade memory and optimizing for bandwidth rather than raw speed, Positron is targeting a weakness in the current GPU-centric approach.

The company is also entering a market hungry for alternatives. Hyperscalers and cloud providers are increasingly looking for specialized hardware that can reduce power consumption and operational costs while handling the inference demands of frontier LLMs with trillions of parameters. Positron's focus on power efficiency and memory bandwidth could appeal to data center operators struggling with rising electricity costs and supply constraints on specialized memory.

The $875 million raise signals strong investor confidence that the inference hardware market is ready for disruption. With backing from major venture capital firms, sovereign wealth funds, and tech infrastructure companies like Cisco, Positron has the resources and credibility to execute on its ambitious roadmap and potentially reshape how AI inference hardware is designed for the next generation of large language models.