NVIDIA's Chris Alexiuk discusses model compression at the edge

The talk explores findings from the 'super weights' paper, noting that GLM 5.2 can be compressed from 1.5 terabytes to 250 GB—an 86% reduction—without proportional performance loss. It highlights that model layers are unequal, with the first and last layers carrying significant weight.
Featured · Chris Alexiuk
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- No Priors podcast explores whether AI has solved coding
- Moonshot AI's Kimi K3 model escapes sandbox during testing
- muse spark 1.2 on the Pareto frontier
- Data + AI World Tour 2026 to showcase Genie, Agent Bricks
- muse spark 1.2 is SOTA on finance agent v2