AnalysisAI ModelsJuly 21, 2026

Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory

Benchmarks report real token-per-second numbers for GPT-OSS 120B, Qwen3.6 MoE, and Hermes agents running locally on AMD Ryzen AI Max's 128GB unified memory. Tests cover speculative decoding and LM Studio on the hardware.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory — AIBriefs