User shares llama.cpp settings for Qwen 3.6 27B on RTX 5090
r/LocalLLaMA poster runs Qwen 3.6 27B on a 5090 at roughly 80-100 t/s for most tasks, dropping to ~40 t/s at the full 262k context. The model barely fits in the 5090's VRAM with vision disabled.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Anthropic's Ultracode coding mode gains industry attention
- Testing the Motion Context node for Stable Diffusion
- Krea2 Turbo BBOX fine-tune uploaded to HuggingFace
- Reddit user shares Minimax H3 character/object V2V swapping template
- ChatGPT accidentally makes photorealistic image mistaken for real photo