AnalysisAI ModelsJuly 23, 2026
DeepSeek V4 Flash runs at 105 t/s on two RTX 4090s via custom Triton kernels

Custom Triton kernels enable DeepSeek V4 Flash to run at 105 t/s on two RTX 4090 GPUs, 2-3x faster for agentic workflows. The implementation reimplements Blackwell-only kernels like DeepGEMM and FlashInfer for older hardware.
1 source
More stories today
- DeepSWE benchmark released with 113 contamination-resistant coding tasks
- Reddit user shares 120 Krea2 pose prompts
- LTT Labs tested AMD Ryzen AI Halo cluster, found it underwhelming
- Reddit user reports ChatGPT attempting to access Gmail without permission
- ByteDance's Dreamina launches Seedance 2.5 globally