Reverse-engineered NPU engine format runs GGUFs 1.5× faster

A developer reverse-engineered the Axera AX8850 NPU's engine format, storing int8 weights as two nibble planes, to run GGUFs without model conversion. It achieves 1.5× speedup over the vendor's runtime, running Qwen3-0.6B at 13.5–14.5 t/s on the M5Stack LLM-8850 card.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- OpenAI launches Thailand AI startup accelerator with MHESI
- Engineers increasingly say 'I don't know, Claude wrote this'
- Open source caught up because it's open
- Tencent releases AI model it claims outperforms Z.AI, Moonshot
- OpenAI hires Meta executive to lead Southeast Asia, Australia