Nvidia says its Groq 3 LPX racks hit 3,400 tokens per second in an Artificial Analysis benchmark on Gemma 4 31B with a 100,000‑token input.
benchmarks gemma gpu inference nvidia
| Source: Techmeme | Original article
Nvidia reports its Groq 3 LPX racks achieved 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000‑token input.
Nvidia announced that its Groq 3 LPX inference racks achieved a throughput of 3,400 tokens per second on the Artificial Analysis benchmark, running the Gemma 4 31B model with a 100,000‑token context. The figure, disclosed in a statement to The Register, demonstrates the accelerator’s ability to sustain high‑speed generation across very long input sequences, a capability Nvidia positions as essential for “highly responsive agentic systems.”
The performance claim follows Nvidia’s earlier announcement that the Groq 3 LPX is now in full production. The platform, built on the Vera Rubin NVL72 infrastructure, aggregates 256 language‑processing units per rack and leverages deterministic, compiler‑scheduled workload planning to avoid the latency spikes that can plague conventional GPU‑based inference. By delivering more than 3 k tokens per second without sacrificing precision, the system aims to close the gap between large‑language‑model reasoning and real‑time interaction, a bottleneck for applications such as conversational assistants, autonomous agents and long‑context analysis tools.
Why it matters is twofold. First, the benchmark underscores Nvidia’s $20 billion investment in Groq’s LPU technology, suggesting the acquisition is beginning to pay off in tangible performance gains. Second, the result puts Nvidia in direct competition with other inference‑focused silicon, where speed at extended context lengths has become a differentiator for enterprise AI workloads.
What to watch next: Nvidia will likely showcase the Groq 3 LPX in customer deployments, building on the Nebius contract we reported on 24 August. Further benchmark releases—especially against rival accelerators—will clarify whether the claimed token rate translates into cost‑effective scaling for production workloads. Observers will also be keen to see how the platform integrates with upcoming model releases that push context windows even farther, potentially cementing Nvidia’s role in the next wave of interactive AI.
Sources
Back to AIPULSEN