AnalysisAI ModelsJuly 23, 2026
Developer ports DeepSeek V4 kernels to Triton for Nvidia 4090d

The implementation enables DeepSeek V4 Flash to run at ~105 tokens/s on dual Nvidia 4090d GPUs by porting Blackwell-specific kernels like DeepGEMM and sparse-MLA to Triton. This achieves a 2-3x performance increase for parallel agentic workflows on sm89 hardware.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Thoughtworks' Kief Morris: humans must stay 'on the loop' in AI delivery
- GEMA wins major copyright ruling against Suno, orders damages paid
- LangChain builds ReviewBench benchmark for code review agents
- DeepSeek Flash 0731's reasoning trace amuses with 'OH MY GOD' outburst
- Former OpenAI VP Jerry Tworek discusses AI lab automation