AnalysisAI ModelsAugust 29, 2026

Qwen 3.8 27B runs at 50 tok/s with 100k context on 16GB GPU

A Reddit user shares a setup running Qwen3.8-27B at 50 tok/s with 100k context on an RTX 4070 Ti SUPER (16GB VRAM) using beellama.cpp and an IQ4_XS GGUF quantization.

1 source

More stories today

Open the live feed