Logo
FrontierNews.ai

NVIDIA's Groq 3 LPX Enters Production: What 3,400 Tokens Per Second Means for AI Agents

NVIDIA has launched Groq 3 LPX, a specialized inference accelerator now in full production that generates AI responses at record speed: 3,400 output tokens per second on benchmark tests. The technology extends NVIDIA's Vera Rubin platform and targets agentic AI systems, which are AI agents that can reason, act, and complete multi-step tasks autonomously. For context, a token is roughly equivalent to a word or small piece of text, so this speed means the system can produce roughly 3,400 words of AI-generated output every second.

Why Does Token Generation Speed Matter for AI Agents?

Agentic AI systems work differently from the chatbots most people interact with. Instead of generating one response and stopping, agents loop through multiple steps: they might read files, write code, call tools, check results, and iterate. Each loop requires the AI to generate tokens, and slow token generation creates bottlenecks. Groq 3 LPX is purpose-built to eliminate these delays, enabling agents to complete tasks in minutes rather than hours.

NVIDIA reports that Groq 3 LPX delivers up to 4 times faster responsiveness compared to the nearest competing platform for latency-sensitive workloads. The accelerator achieved its 3,400 tokens-per-second benchmark while processing the Gemma 4 31B model, an open-source AI model designed for agentic tasks, with a 100,000-token context window, meaning the system could reference roughly 100,000 words of prior conversation or documents simultaneously.

How Does Groq 3 LPX Fit Into NVIDIA's Broader AI Strategy?

Groq 3 LPX is not a standalone product. It extends the Vera Rubin NVL72 platform, which NVIDIA positions as a comprehensive "AI factory" architecture. The Vera Rubin ecosystem includes multiple specialized components designed to handle different AI workloads, from training massive models to running inference at scale. Groq 3 LPX specifically optimizes the inference phase, the stage where trained models generate outputs for users.

The platform integrates with other NVIDIA infrastructure components including BlueField-4 data processing units (DPUs), which handle networking and security, and Spectrum-6 Ethernet for high-speed data transfer. This extreme codesign across seven chips and five purpose-built racks reflects NVIDIA's strategy of building tightly integrated systems rather than selling individual components.

"Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency. Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation," said Jensen Huang, founder and CEO of NVIDIA.

Jensen Huang, Founder and CEO at NVIDIA

Which Companies Are Adopting Groq 3 LPX First?

Nebius, a leading AI cloud provider, will be the first to deploy Groq 3 LPX in production through its Nebius Token Factory platform, which serves developers building inference-heavy applications. Groq, a purpose-built AI inference cloud company, plans to be among the earliest additional adopters. Both companies target the growing market of enterprises and developers who need fast, responsive AI inference at scale.

"Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate. As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant," stated Danila Shtan, chief technology officer of Nebius.

Danila Shtan, Chief Technology Officer at Nebius

What Are the Key Technical Advantages of Groq 3 LPX?

  • Record Token Generation Speed: Achieved 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context, the fastest performance ever recorded for this model in independent benchmarking.
  • Responsiveness for Agentic Workloads: Provides up to 4 times faster responsiveness than competing platforms, enabling AI agents to complete complex reasoning and coding tasks in minutes instead of hours.
  • Context Window Support: Handles 100,000-token contexts, allowing agents to reference vast amounts of prior information, documents, or conversation history while maintaining speed.
  • Integration with Vera Rubin Ecosystem: Works seamlessly with BlueField-4 DPUs, Vera CPU racks, and Spectrum-6 Ethernet as part of a comprehensive AI factory architecture.

How Can AI Cloud Providers Leverage Groq 3 LPX?

For AI cloud providers serving enterprises and developers, Groq 3 LPX offers a path to differentiate their offerings. Latency-sensitive, high-volume inference workloads, such as real-time coding assistance, autonomous research agents, and interactive AI applications, benefit most from the accelerator's speed. Nebius plans to expose Groq 3 LPX capabilities through the same API developers already use, eliminating the need for migration to new tools or workflows.

The timing of Groq 3 LPX's production launch reflects broader industry momentum. AI cloud providers are becoming central to the AI economy, offering enterprises and developers access to advanced infrastructure for training, reasoning, and inference at scale. As demand for AI computation accelerates worldwide, the ability to deliver fast, responsive inference becomes a competitive advantage.

NVIDIA's announcement of Groq 3 LPX in full production signals that the company views agentic AI as a major growth opportunity. Unlike traditional large language models that generate a single response, agentic systems require repeated, fast token generation across multiple reasoning steps. By optimizing for this workload pattern, NVIDIA is positioning itself at the center of the next phase of AI infrastructure, where speed and responsiveness define user experience.