NVIDIA Groq 3 LPX enters full production, hits 3,431 tokens/s on Gemma 4 31B

NVIDIA announced Groq 3 LPX, an interactive inference accelerator for Vera Rubin, is in full production. Artificial Analysis measured 3,431 output tokens/sec on Gemma 4 31B with 100K context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.
How this story unfolded
same day · 3 reports · 2 community posts · 5 of 6 shown
- Aug 24
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated...
- Aug 25
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Stable Diffusion user tests H3 model with Cheers-style script
- Reddit users share impressive image-to-video AI demos
- Reddit reminds users they can legally seed AI models via torrenting
- MiniMax H3 reverse-engineers paintings into basic forms
- OpenAI DevDay Exchange Seoul applications close Sept 4