Developer builds 250M-parameter quantized LLM that runs in 60 MB

A developer trained a 250M-parameter model from scratch on 30B tokens of fineweb, quantized to under 2 bits, deploying in 60 MB and running at ~400 tok/s on a laptop CPU with no GPU.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude detects failing drive, saves user's data
- AI and satellite guidance could bring robot mowers to half of US lawns
- OpenAI acquires Instant backend team
- China unveils AI-powered flying lifebuoy for water rescues
- Apollo's Slok: AI weighs on pay without cutting jobs