AnalysisDevelopersAugust 5, 2026

Mference runs Inkling-Small 276B-A12B on <10GB at ~2.9 tok/s

Mference now supports Thinking Machines' Inkling-Small 276B-A12B (Apache 2.0) via a 4-bit MLX conversion, delivering ~2.9 tok/s on under 10GB of memory. Runs from the pipenetwork/Inkling-Small-MLX-4bit repo.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed