LFM2.5-2.6B runs at 30 tok/s on a phone

Liquid AI's LFM2.5-2.6B, a 2.69B-parameter model with 128K context and tool calling, runs at 30 tok/s on a phone via Q4_K_M GGUF. A custom engine achieves 17 tok/s on a OnePlus 13.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GPT Images viral again on Reddit
- Suno users rush to generate 4.5+ Pro songs before deadline
- Qwen 3.8 35B A3B requested by users for speed
- Why do most tech subs seem to hate Claude so much?
- Tutorial: Build document intelligence pipeline with deepDoctection