AnalysisAI ModelsJuly 20, 2026
1-bit quant of Hy3 295B runs 2.2x faster than cloud API without quality loss

Community quantization of Tencent's Hy3 295B model to 1-bit produces a 92GB IQ1_M GGUF file that runs locally on 4x RTX 5090. In tests, the quantized model matched the cloud API's quality on a retro game generation task while running 2.2x faster. The result suggests extreme quantization can preserve capability for some workloads.
1 source
More stories today
- America’s Open-Model Paradox
- GEMA launches fully cleared PLAI music dataset for AI developers
- Claude Code 2.1.220 release imminent
- Gary Marcus writes open letter to David Sacks on AI regulation
- Prentis AI lab co-founded by Hoffman, Pincus in talks to raise $100M