DeepSeek V4 Flash: CUDA 13.1 yields 100-150 faster t/s in prefill

A r/LocalLLaMA post reports DeepSeek V4 Flash runs 100-150 faster t/s in prefill with CUDA 13.1 over 13.3. CUDA 13.2 is said to be buggy; a vibed fork works with CUDA 13.3 as the alternative.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs