AnalysisDevelopersJune 30, 2026
NVIDIA's software stack cuts DeepSeek V4 token costs by 5x on Blackwell

NVIDIA's inference software stack on Blackwell GPUs reduces token costs by up to 5x on DeepSeek V4 within a month. Customers like Baseten, Cognition, and Together AI are leveraging TensorRT-LLM and Dynamo for production inference. DigitalOcean helped Hippocratic AI achieve 30% higher throughput on Blackwell.