NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA announced Groq 3 LPX, a token-generation accelerator for Vera Rubin NVL72, is in full production. In an Artificial Analysis benchmark on Gemma 4 31B, it delivered 3,400 output tokens per second at 100K context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.

How this story unfolded

3 days · 11 reports · 2 community posts · 13 of 16 shown

  1. Aug 24
  2. Aug 26
  3. Aug 27

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed