NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA announced Groq 3 LPX, an interactive inference accelerator for Vera Rubin, is in full production. In an Artificial Analysis benchmark on Gemma 4 31B, it delivered 3,400 output tokens per second at 100K context, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.

How this story unfolded

3 days · 8 reports · 1 community post · from Aug 24

  1. Aug 24
  2. Aug 26
  3. Aug 27

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed