NVIDIA on August 24, 2026 said its Groq 3 LPX interactive AI inference accelerator has entered full production, extending the Vera Rubin platform with world-class speed for agentic AI.
An extension of the NVIDIA Vera Rubin architecture, Groq 3 LPX is designed to generate tokens at ultrafast rates for interactive agents that must respond in real time. NVIDIA said the chip works in concert with Spectrum-X networking and NVLink Fusion to let every layer of the "AI factory" operate together.
"The next era of AI inference won't be defined by a single breakthrough chip, network or system," NVIDIA wrote. "It'll be defined by how every layer of the AI factory works together." The company positioned Groq 3 LPX as the interactive-inference complement to its training-class accelerators.
The production milestone signals that NVIDIA's Rubin-generation lineup is now shipping across both training and low-latency inference roles, tightening the integration of its full-stack AI platform.





