NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering 3,431 output tokens per second on Gemma 4 31B with 100K context in Artificial Analysis benchmark, 4x faster than nearest alternative. Nebius is first AI cloud to adopt it.
How this story unfolded
12 days · 12 reports · 1 community post · 13 of 16 shown
- Aug 12
- Aug 22
- Aug 24
How XPUs Meet a World-Class AI Factoryblogs.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scalenvidianews.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rulesdeveloper.nvidia.com
NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Wattdeveloper.nvidia.com
NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factoriesdeveloper.nvidia.com
Solving Agentic AI Fleet Challenges with NVIDIA Vera CPUdeveloper.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agentsblogs.nvidia.com
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated...
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Robotic systems struggle with physical navigation in recent demonstrations
- ChatGPT Visualize skill turns notes into interactive interfaces
- Twitch streamers sue Twitch, Amazon over AI training data use
- Fastino launches GLiNER2.5 with boundary-prediction architecture
- NVIDIA Sol Engine cuts MiniMax H3 latency to 14.93s