SANTA CLARA — Nvidia announced on August 24 the full mass production ramp of its Groq 3 LPX AI inference chip, alongside new performance benchmarks for the Vera Rubin platform that the company says will redefine the economics of AI infrastructure in the a
SANTA CLARA — Nvidia announced on August 24 the full mass production ramp of its Groq 3 LPX AI inference chip, alongside new performance benchmarks for the Vera Rubin platform that the company says will redefine the economics of AI infrastructure in the agentic computing era.
In SemiAnalysis AgentX agentic workload testing running the DeepSeek V4 Pro model, the Vera Rubin NVL72 system achieved up to 30 times the per-megawatt throughput of its predecessor, the GB300 NVL72, while reducing per-token cost by up to 35 times. The Groq 3 LPX chip, focused on low-latency token generation, achieved 3,400 output tokens per second on the Gemma 4 31B model across a 100,000-token context — a benchmark that positions it for real-time inference applications requiring rapid response.
SpaceX AI announced it would expand deployment of Nvidia's newest-generation platforms, adopting Vera Rubin CPUs for next-generation agentic AI applications and building Vera Rubin NVL72 infrastructure as its own compute capacity scales toward multiple gigawatts. Most significantly, SpaceX AI plans to deploy optimized Vera Rubin NVL72 systems on its first generation of Starmind AI satellites, extending Nvidia's AI computing platform from ground-based data centers into orbital compute — a concept that, if commercially viable, would represent a fundamental expansion of where AI inference can take place.
Nvidia's dual product strategy — the GB300 NVL72 as the training backbone and Vera Rubin as the inference-optimized successor — reflects the industry shift toward larger-scale inference workloads as AI deployments move from training to production. With AI agents increasingly deployed to autonomously execute multi-step tasks, the volume of inference compute per query is growing far faster than training compute, making inference efficiency a primary architectural concern for hyperscalers and enterprise AI deployers alike.