AnalysisAI ModelsAugust 10, 2026

Tiny LLM runs at 59,965 tok/s entirely on-chip on a $250 FPGA

A 3.16M-parameter INT4 transformer runs entirely in the on-chip memory of a $250 Xilinx Kria KV260 FPGA, hitting 59,965 tok/s with zero DRAM in the token loop. The same model manages 11 tok/s on the board's own Arm cores and 719 tok/s on a laptop RTX 3050 Ti.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed