AnalysisAI ModelsJuly 23, 2026

DeepSeek V4 Flash runs at 105 t/s on two RTX 4090s via custom Triton kernels

Custom Triton kernels enable DeepSeek V4 Flash to run at 105 t/s on two RTX 4090 GPUs, 2-3x faster for agentic workflows. The implementation reimplements Blackwell-only kernels like DeepGEMM and FlashInfer for older hardware.

1 source

More stories today

Open the live feed