NVIDIALaunchDevelopersSeptember 15, 2026

NVIDIA Groq 3 LPX enters full production for agentic inference

NVIDIA's Groq 3 LPX rack-scale accelerator is now in full production, extending Vera Rubin NVL72 with fast token generation for agentic systems. Artificial Analysis measured 3,431 tokens/s on Gemma 4 31B at 100k context, the highest it has recorded for that model on any endpoint.

How this story unfolded

3 weeks · 4 reports · 2 community posts · 6 of 7 shown

  1. Aug 24
  2. Aug 25
  3. Sep 15

More stories today

Open the live feed