AnalysisDevelopersJune 30, 2026
NVIDIA's inference software stack reduces token costs by up to 5x on Blackwell

NVIDIA reports up to 5x reduction in token costs for DeepSeek V4 on Blackwell GPUs within one month, driven by its full-stack inference software including TensorRT-LLM and Dynamo. Companies like Baseten, Cognition, Deep Infra, DigitalOcean, and Together AI are already leveraging the optimizations for production workloads.