DeepSeek V4 Flash on 5090 in llama.cpp with 1 Million context

User shared benchmarks and config for running Unsloth's DeepSeek-V4-Flash-UD-Q8_K_XL quant in llama.cpp on an RTX 5090 with 1M context, saying recent llama.cpp changes made the model much more usable.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- CoreWeave to Enter Asian Market With Indonesian Data Centers
- Nvidia, Dell Back AI Cloud Startup Volta at $2.4 Billion Value
- ESPN unveils AI tells detection at World Series of Poker
- Gemini Agent-to-Agent Attack Exposed Secrets, Enabled Pull Request Tampering
- Podcast examines Hollywood's quiet embrace of AI and control battle