AnalysisAI ModelsSeptember 3, 2026

Qwen3.8-27B runs at 130 tok/s on Kaggle's free TPU

A user got Qwen3.8-27B in bf16 with no quantization running on Kaggle's free TPU v5e-8 at ~130 tok/s single-stream and 10,300 tok/s prefill, with the full native 262,144-token context. Other local setups report ~280 tok/s on dual Radeon AI Pro R9700s and ~75 tok/s decode on a 16GB RTX 5080.

1 source

More stories today

Open the live feed