NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA announced Groq 3 LPX, an interactive inference accelerator for Vera Rubin, is in full production. In an Artificial Analysis benchmark on Gemma 4 31B, it delivered 3,400 output tokens per second at 100K context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.
How this story unfolded
3 days · 8 reports · 1 community post · from Aug 24
- Aug 24
How XPUs Meet a World-Class AI Factoryblogs.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated...
- Aug 26
- Aug 27
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Open-source RL training with trl and OpenEnv shared
- NVIDIA invests $3.5B in MediaTek, deepens AI partnership
- Neta team explains why their open-source model generated Anne Hathaway-like images
- OpenAI age-verification error deletes adult's account
- South Korea gives citizens free unlimited domestic AI access