Run an on-chip language model at 60k tok/s on a $250 FPGA
This is a working demonstration of a 3.16M-parameter INT4 transformer running entirely inside the on-chip memory of a $250 Xilinx Kria KV260 FPGA. The model — a TinyStories-class story generator at roughly 1.5 MB of weights — lives in URAM/BRAM with zero DRAM in the token loop, achieving a bit-exact 59,965 tokens per second measured on silicon. A live WebSocket demo on the page connects you to the actual board in Wales so you can chat with it in real time. It is a story generator, not an assistant: give it 'once upon a time' and it finishes the story in deliberately telegraphic style, because the compression is the speed. The project is aimed at FPGA developers, ML engineers, and edge-AI hobbyists who want to see what can be squeezed into on-chip fabric memory without touching the shared DDR bandwidth wall. It also serves as a practical study of memory-bound transformer inference on a low-cost Zynq UltraScale+ board.