NVIDIA Groq 3 LPX Inference Accelerator Enters Full Production
en.Wedoany.com Reported - On August 24, NVIDIA announced at the Hot Chips 2026 conference that its Groq 3 LPX accelerator, designed specifically for AI agent inference, has entered full production, marking the commercial realization of the company's first product following its $20 billion acquisition of Groq chip technology licensing in December 2025.

The Groq 3 LPX is a specialized inference accelerator for NVIDIA's Vera Rubin data platform, focusing on the "decode" phase of AI inference—the critical stage that determines token generation rate. A single LPX rack integrates 256 Groq 3 LPU accelerators, equipped with 128 GB of on-chip SRAM and 640 TB/s of inter-chip scaling bandwidth, adopting NVIDIA's MGX rack architecture with full liquid cooling. The chip is manufactured by Samsung Electronics, creating a differentiated supply chain layout compared to NVIDIA GPUs (manufactured by TSMC).
In benchmark tests conducted by third-party organization Artificial Analysis, the Groq 3 LPX achieved a generation speed of 3,400 output tokens per second when running the open-source model Gemma 4 31B with a 100K token long context. NVIDIA claims this is the fastest inference performance ever recorded for this model, delivering 4x faster performance on latency-sensitive tasks compared to competing platforms.
AI cloud service provider Nebius has signed on as the first deployment customer for the Groq 3 LPX, offering developers access to the accelerator's ultra-fast token generation capabilities through the Nebius Token Factory inference platform. Dion Harris, Senior Director at NVIDIA, stated that Groq racks will be deployed alongside Vera central processors and Rubin graphics processors at Nebius, with full availability expected in the second half of 2026. Nebius CTO Danila Shtan remarked: "Generation is the inference phase that determines AI system response speed, and that is exactly what Groq 3 LPX accelerates."
NVIDIA founder and CEO Jensen Huang stated that the Vera Rubin platform, through the Groq 3 LPX, has achieved an optimized compute factory configuration for the AI agent era, "transforming the way intelligence is produced, delivering another massive leap in AI throughput, efficiency, and response speed".
Harris emphasized that the Groq 3 LPX is not intended to replace GPUs but rather to use the optimal processor for specific workloads—GPUs continue to handle model training and general inference, while Groq chips focus on the latency-sensitive "decode" phase. NVIDIA also announced that SpaceX has signed on as the latest flagship customer for the Vera Rubin platform, deploying Vera CPUs for CPU-intensive orchestration and simulation tasks spanning ground data centers and orbital satellites. Shipments of the Vera Rubin system are accelerating, with production having commenced in early 2026.
Related Products




Hot Selling 10 Inch Intel N5100 I5 I7 FHD 8 +128GB 16+256GB Win10 Pro Win11 Pro 4G 5G Rugged Tablet PC Pad with NFC Barcode
Highton Electronics Co., Ltd.

Comprehensive Mining Automation Control System
Beijing Tianma Intelligent Control Technology Co., Ltd.

Automatic Aiming Laser Remote Obstacle Removal Robot
Pinggao Group Weihai High-Voltage Apparatus Co., Ltd.
MUX Series Communication Interface Device for Security and Stability Control System
Nanjing NR Electric Co., Ltd.
20 Inch 1600*900 Full HD 60Hz TFT LED Monitor with VGA and DP Interfaces for Desktop Computers-Business Use
Guangzhou Zanying Optoelectronics Technology Co., Ltd.










