NVIDIA Vera Rubin NVL72 delivers 30x more work per watt for AI agents

NVIDIA's Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads, with 35x lower token cost, per SemiAnalysis AgentX benchmark. Agentic requests consume 15x more tokens than chat. Microsoft has first operational racks; SpaceXAI adopts Vera CPU.
How this story unfolded
4 days · 10 reports · 3 community posts · 13 of 15 shown
- Aug 22
- Aug 24
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agentsblogs.nvidia.com
SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scalenvidianews.nvidia.com
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AInvidianews.nvidia.com
NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Wattdeveloper.nvidia.com
Solving Agentic AI Fleet Challenges with NVIDIA Vera CPUdeveloper.nvidia.com
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubindeveloper.nvidia.com
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com
- Aug 25
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Prompt search tool indexes 10,000+ image generation prompts
- Judge refuses to toss indie musician's lawsuit against Suno
- Nvidia sells first H200 chips in China, shipments below allowed total
- Model company acquisitions discussed
- Apple's Luce generates relightable 3D assets from single images