NVIDIA Groq 3 LPX enters full production for Vera Rubin

NVIDIA announced Groq 3 LPX, its interactive AI inference accelerator for the Vera Rubin platform, is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, it delivered 3,431 output tokens per second on 100K context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.
3 sources
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Meta builds AI 'second brain' that learns from experts
- Snowflake shares surge 22% on AI coding momentum
- Gary Marcus critiques Musk's AI prediction shift
- Lutnick says Anthropic has patched relations with US government
- HPE lifts sales forecast on AI server demand