AnalysisDevelopersSeptember 8, 2026

Custom llama.cpp build targets 7900 XTX with PCIe x4 and tensor parallel tweaks

A community build for AMD 7900 XTX cards reports 920 tok/s prompt processing at 8192 tokens on Qwen 3.8 Next Q3_K_XL with two cards and RAM offload, plus 24-27 tok/s on prose and 40+ tok/s on code with MTP.

1 source

More stories today

Open the live feed