Mference runs Inkling-Small 276B-A12B on <10GB at ~2.9 tok/s

Mference now supports Thinking Machines' Inkling-Small 276B-A12B (Apache 2.0) via a 4-bit MLX conversion, delivering ~2.9 tok/s on under 10GB of memory. Runs from the pipenetwork/Inkling-Small-MLX-4bit repo.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Daydream launches AI companion virtual friend app
- EU AI Act prompts text watermarking; detection API to ship
- OpenAI launches ChatGPT and Codex desktop app for Linux
- Landscape map charts path of self-evolving AI agents
- Zitron: 70% of hyperscaler AI revenue flows from OpenAI, Anthropic