AnalysisAI ModelsJuly 21, 2026
Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory

Benchmarks report real token-per-second numbers for GPT-OSS 120B, Qwen3.6 MoE, and Hermes agents running locally on AMD Ryzen AI Max's 128GB unified memory. Tests cover speculative decoding and LM Studio on the hardware.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Thoughtworks' Kief Morris: humans must stay 'on the loop' in AI delivery
- GEMA wins major copyright ruling against Suno, orders damages paid
- LangChain builds ReviewBench benchmark for code review agents
- DeepSeek Flash 0731's reasoning trace amuses with 'OH MY GOD' outburst
- Former OpenAI VP Jerry Tworek discusses AI lab automation