NVIDIA Groq 3 LPX enters mass production, Vera Rubin inference 30x prior gen
NVIDIA announced Groq 3 LPX mass production on Aug 24; Vera Rubin NVL72 achieves 30x GB300 throughput on AgentX workload.
On August 24, NVIDIA announced that Groq 3 LPX, its interactive AI inference accelerator, has entered mass production. As part of the Vera Rubin platform, Groq 3 LPX focuses on ultra-low-latency token generation. In Artificial Analysis testing with the Gemma 4 31B model and 100,000-token context, it reached 3,400 output tokens per second, a new record for that model.
NVIDIA also released latest test results for Vera Rubin NVL72 on SemiAnalysis's AgentX workload running the DeepSeek V4 Pro model: per-megawatt throughput reaches up to 30x the previous-generation GB300 NVL72, with per-token cost reduced by up to 35x.
On the same day, SpaceXAI announced it will adopt Vera CPU to accelerate next-generation agent AI applications, and will extend an optimized Vera Rubin NVL72 to space aboard first-generation Starmind AI satellites. xAI had previously confirmed its AI infrastructure will fully migrate to the Vera Rubin platform.