NVIDIALaunchAI ModelsAugust 17, 2026

NVIDIA releases Nemotron 3.5 Lightning, open 30B MoE agentic model

Nemotron 3.5 Lightning pairs 30B LatentMoE parameters with 3B active, delivering up to 4x output speed and up to 1M context. The NVFP4 checkpoint compresses 66 GB to 22 GB; Perplexity Agent API pricing is $0.0115 input / $0.17 output per 1M tokens. Available open on Hugging Face and SageMaker JumpStart.

How this story unfolded

3 weeks · 24 reports · 14 community posts · 38 of 44 shown

  1. Jul 28
  2. Aug 4
  3. Aug 11
  4. Aug 12
  5. Aug 13
  6. Aug 17
  7. Aug 18

More stories today

Open the live feed