NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/s

NVIDIA announced Groq 3 LPX, an interactive AI inference accelerator for Vera Rubin NVL72, is now in full production. In Artificial Analysis benchmarking running Gemma 4 31B, it delivered 3,400 output tokens per second with a 100,000-token context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.
Featured · Jensen Huang
How this story unfolded
same day · 3 reports · 2 community posts · 5 of 6 shown
- Aug 24
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated...
- Aug 25
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Prompt search tool indexes 10,000+ image generation prompts
- Judge refuses to toss indie musician's lawsuit against Suno
- Nvidia sells first H200 chips in China, shipments below allowed total
- Model company acquisitions discussed
- Apple's Luce generates relightable 3D assets from single images