Tiny LLM runs at 59,965 tok/s entirely on-chip on a $250 FPGA

A 3.16M-parameter INT4 transformer runs entirely in the on-chip memory of a $250 Xilinx Kria KV260 FPGA, hitting 59,965 tok/s with zero DRAM in the token loop. The same model manages 11 tok/s on the board's own Arm cores and 719 tok/s on a laptop RTX 3050 Ti.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Code 2.1.227 fixes subscription-tier, Bash and TUI bugs
- Curated resources for the open Agent2Agent protocol
- Suno to cap song downloads to curb AI slop
- Claude Code plugin translates 'Claudish' output into plain English
- Claude Code v2.1.227 fixes flag evaluation and Bash command failures