How-ToAI ModelsJuly 15, 2026

Gemma 4 26B runs at 5 tokens/sec on 13-year-old Xeon without GPU

A 13-year-old Xeon CPU achieves 5 tokens/sec inference with Gemma 4 26B via aggressive quantization and memory tuning. The setup uses 4-bit quantization and custom kernel optimizations, demonstrating viability of large model inference on legacy hardware.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
Gemma 4 26B runs at 5 tokens/sec on 13-year-old Xeon without GPU