DeepSeek-V4-Flash-0731 Dwarfstar hits 28 t/s decode on M2 Ultra

On an M2 Ultra with 192GB of RAM, decode starts at 28 tokens/s, holding at 23.5 t/s at 45k context and 18 t/s at 192k depth with 8k-token output maintained.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs