NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering 3,431 output tokens per second on Gemma 4 31B with 100K context in Artificial Analysis benchmark, 4x faster than nearest alternative. Nebius is first AI cloud to adopt it.

How this story unfolded

12 days · 12 reports · 1 community post · 13 of 16 shown

  1. Aug 12
  2. Aug 22
  3. Aug 24

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed