DeepSeek V4 Flash hits 32 tok/s on AMD Ryzen AI MAX+ 395

DeepSeek V4 Flash plus its speculative draft fits on a single AMD Ryzen AI MAX+ 395 with 128 GB unified memory, reaching a usable decode rate of up to 32 tok/s, per a r/LocalLLaMA post from the team behind the work.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- NVIDIA video explains how agentic AI powers autonomous networks
- SoftBank Earnings to Test Appetite for AI Bets Beyond ChatGPT
- Pixel-Native RAG: A Practical Guide to Visual Document Indexing
- User demonstrates Minimax video generation workflow with SageAttention
- User faces backlash after using Claude to process audio from toddler sleepover