AnalysisDevelopersJuly 25, 2026
LFM 2.5 230M running at 1440 tok/s in-browser through a custom backend

A custom WebGPU backend achieves 1440 tokens/s for LFM 2.5 230M entirely in-browser. It supports Nvidia with fused multi-pass kernels and Apple Silicon via Metal, running in browser or Electron/Tauri apps.
1 source
More stories today
- Open-source model announced by Hugging Face
- Smartphone recording powers fast, cheap motion capture alternative
- User uses Claude to reduce medical bill from $1,200 to $180
- ai-sage releases GigaChat3.1-Audio-10B audio LLM
- Reddit user runs six-week experiment on faceless AI persona accounts