AnalysisDevelopersAugust 6, 2026

Mference runs Inkling-Small 276B-A12B on <10GB at ~2.9 tok/s

Read original source →reddit.com

Mference now supports Thinking Machines' Inkling-Small 276B-A12B (Apache 2.0) via a 4-bit MLX conversion, delivering ~2.9 tok/s on under 10GB of memory. Runs from the pipenetwork/Inkling-Small-MLX-4bit repo.

1 source

More stories today

Open the live feed
Mference runs Inkling-Small 276B-A12B on <10GB at ~2.9 tok/s