How Cactus Bonsai runs a 27B-parameter model on a phone

Cactus Bonsai compresses a 27-billion-parameter LLM to 3.9GB using 1-bit quantization and quantization-aware training, letting it run locally on a smartphone. A standard-precision version would need over 50GB of memory — roughly 108GB at FP32 precision.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Lap: open-source photo manager runs face recognition locally
- MemGraphRAG is a memory-based multi-agent system for Graph RAG
- MD-This-Page converts any webpage into LLM-ready Markdown
- NVIDIA releases Molt, a PyTorch-native agentic RL framework
- Superior Skills offers open-source agent trading schemas on Hyperliquid