Qwen 3.6 27B flags and performance in llama.cpp on RTX 5090
User reports Qwen 3.6 27B runs at 80-100 t/s on an RTX 5090 via llama.cpp, dropping to ~40 t/s at full 262k context. The setup barely fits in VRAM without vision and is used primarily for appdev tasks.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs