AnalysisAI ModelsAugust 16, 2026

Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM

A user tested Qwen3.8 27B dense at Q2 and Q3 against Qwen3.6 35B-A3B MoE on a 12GB RTX 5070 Ti Laptop using llama.cpp CUDA with 4k context and q8 KV. Results show which quantizations are usable on 12GB VRAM.

1 source

More stories today

Open the live feed