Shoehorn quantizes models to fit your machine's memory

Open-source tool starts from your actual memory budget and solves a per-tensor mixed-precision assignment, routinely using 99.99% of it. Uses llama.cpp as the backend, with a browser-based GUI that scans Hugging Face's most-downloaded models and streams the fit.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs