NVIDIALaunchDevelopersAugust 24, 2026

NVIDIA Groq 3 LPX in full production, hits 3,400 tokens/sec

NVIDIA's Groq 3 LPX inference accelerator is now in full production, delivering 3,431 output tokens per second on Gemma 4 31B with 100K context in Artificial Analysis benchmark, 4x faster than nearest alternative. Nebius is first AI cloud to adopt it.

4 sources

NVIDIA by email

Get an email when NVIDIA has news

No news that day, no email.

More stories today

Open the live feed