AnalysisAI ModelsJuly 3, 2026
Deepseek V4 Flash running on RTX 5090 MoE

User optimized Deepseek V4 Flash on RTX 5090 MoE, reporting TG T/S of 21.3 and PP T/S of 927 after optimization. The configuration uses MoE, no unified KV, and no memory map, with prompt processing ranging from 8k to 65k tokens.
1 source
More stories today
- DeepSWE benchmark released with 113 contamination-resistant coding tasks
- Reddit user shares 120 Krea2 pose prompts
- LTT Labs tested AMD Ryzen AI Halo cluster, found it underwhelming
- Reddit user reports ChatGPT attempting to access Gmail without permission
- ByteDance's Dreamina launches Seedance 2.5 globally