AnalysisAI ModelsJuly 23, 2026

Developer ports DeepSeek V4 kernels to Triton for Nvidia 4090d

The implementation enables DeepSeek V4 Flash to run at ~105 tokens/s on dual Nvidia 4090d GPUs by porting Blackwell-specific kernels like DeepGEMM and sparse-MLA to Triton. This achieves a 2-3x performance increase for parallel agentic workflows on sm89 hardware.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed