DeepSeek V4 Flash runs locally on AMD Ryzen AI MAX+ 395 at up to 32 tok/s

Community build fits DeepSeek V4 Flash plus its speculative draft on a single AMD Ryzen AI MAX+ 395 with 128 GB unified memory, reaching a usable decode rate of up to 32 tokens/s. Details are shared in a linked blog post aimed at Strix Halo owners.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- CoreWeave to Enter Asian Market With Indonesian Data Centers
- Nvidia, Dell Back AI Cloud Startup Volta at $2.4 Billion Value
- ESPN unveils AI tells detection at World Series of Poker
- Gemini Agent-to-Agent Attack Exposed Secrets, Enabled Pull Request Tampering
- Podcast examines Hollywood's quiet embrace of AI and control battle