NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering 3,431 output tokens per second on Gemma 4 31B with 100K context in Artificial Analysis benchmark, 4x faster than nearest alternative. Nebius is first AI cloud to adopt it.
4 sources
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
Inference on steroid: NVIDIA’s Groq 3 LPX is now in full production, adding a dedicated...x.com
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Google AI's EnvHarness turns static agent benchmarks into adaptive training worlds
- AI agents that pass authentication can still drift, expose data, or get memory-poisoned
- Local video watermark remover released on CivitAI
- Video walks through implementing Kimi K3 from scratch in PyTorch
- SeedVR2 TensorRT Studio offers free open-source video upscaling