LaunchAI ModelsJuly 16, 2026
Qwen3.5 122B-A10B GGUF with ROCmFP4 iMatrix released

The 122B-parameter model at 60.70 GiB achieves 28.50 tok/s on AMD Strix Halo, 36.89% faster decode and 13.47 GB smaller than comparable quants. Built using the ROCmFP4 format, it requires a custom llama.cpp fork.