How-ToDevelopersAugust 7, 2026

User shares llama.cpp settings for Qwen 3.6 27B on RTX 5090

r/LocalLLaMA poster runs Qwen 3.6 27B on a 5090 at roughly 80-100 t/s for most tasks, dropping to ~40 t/s at the full 262k context. The model barely fits in the 5090's VRAM with vision disabled.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed