NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/s

NVIDIA announced Groq 3 LPX, an interactive AI inference accelerator for Vera Rubin NVL72, is now in full production. In Artificial Analysis benchmarking running Gemma 4 31B, it delivered 3,400 output tokens per second with a 100,000-token context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.

Featured · Jensen Huang

How this story unfolded

same day · 3 reports · 2 community posts · 5 of 6 shown

  1. Aug 24
  2. Aug 25

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed