NVIDIA Groq 3 LPX enters full production for agentic AI

NVIDIA announced Groq 3 LPX, an interactive inference accelerator for Vera Rubin, is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, it delivered 3,400 output tokens per second for 100,000-token contexts, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.
How this story unfolded
3 days · 7 reports · 2 community posts · 9 of 10 shown
- Aug 24
SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scalenvidianews.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
Solving Agentic AI Fleet Challenges with NVIDIA Vera CPUdeveloper.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
- Aug 27
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Redditor tests AI agents with $1 online task
- AI agents need their own identity before a gateway
- Claude Max users find default $200K spend limit
- TTFT-First Benchmark Ranks Lowest-Latency Voice and Realtime Agent APIs
- AI training demand causes Mac Mini shortages