How to run DeepSeek V4 Flash locally with 4-bit quantization

DeepSeek V4 Flash is a 284B-parameter text-only model; local 4-bit deployment needs ~168GB VRAM (~110GB at 3-bit), typically two DGX Spark units. It jumped from ~7% to 54% on the DeepSweep agentic-coding benchmark, and real-world testing hit ~25 tokens/s on a dual DGX Spark cluster.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs