NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering 3,400 output tokens per second on Gemma 4 31B with 100K context in Artificial Analysis benchmarks, 4x faster than the nearest alternative. Nebius is the first AI cloud to adopt it.

How this story unfolded

3 days · 6 reports · 2 community posts · 8 of 10 shown

  1. Aug 22
  2. Aug 24
  3. Aug 25

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed