NVIDIA Groq 3 LPX inference accelerator enters mass production at Hot Chips 2026
On August 25 at Hot Chips 2026, NVIDIA announced Groq 3 LPX mass production with 3,400 output tokens/sec on Gemma 4 31B at 100K context.
On August 25 at the Hot Chips 2026 semiconductor design conference, NVIDIA announced that the Groq 3 LPX inference accelerator has entered mass production. The chip is an extension of the Vera Rubin data-center platform, designed to address the gap traditional GPUs leave in ultra-low-latency token generation.
In an Artificial Analysis benchmark with Gemma 4 31B at 100,000-token context, Groq 3 LPX reached 3,400 output tokens per second, the highest score recorded for the model. For latency-sensitive agentic coding workloads, the chip delivers up to 4x the response speed of recent competing platforms.
On the Vera Rubin NVL72 rack-scale platform, with DeepSeek V4 Pro running the SemiAnalysis AgentX workload, per-megawatt throughput reaches up to 30x the previous-generation GB300 NVL72 and per-token cost falls up to 35x. Nebius Token Factory has committed to be among the first adopters, with Groq joining the early customer list.