AnalysisAI ModelsJuly 21, 2026

Local LLM speed test: GPT-OSS, Qwen3.6, Hermes on 128GB memory

Benchmarks show GPT-OSS 120B achieves X tokens/s, Qwen3.6 MoE Y tokens/s, and Hermes agents Z tokens/s on 128GB unified memory. The hardware is AMD Ryzen AI Max Plus 395 with Radeon 8060S GPU, enabling local 100B+ parameter models without discrete GPU.

2 sources