AnalysisAI ModelsOctober 1, 2026

Community projects squeeze Qwen 3.8 27B onto 16GB GPUs

Read original source →github.com

Ninfer 4080 runs ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB, hitting 2720 tok/s prefill and 262 tok/s generation. A separate guide runs Qwen 3.8 27B Q4 XS at ~30 t/s decode on a 16GB RX 7800 XT via llama.cpp Vulkan.

How this story unfolded

4 days · 0 reports · 3 community posts · from Sep 29

  1. Sep 29
  2. Oct 1
  3. Oct 3

More stories today

Open the live feed