GitHub radar
APEX: Open LLM Inference Chip Design for FPGA
Sigmantic AI has published APEX, an open source, bit-exact-verified hardware implementation of a transformer decoder tile for running real LLM models on FPGA chips. It compresses the KV cache in the datapath to maintain constant read throughput during long conversations.
APEX is an open source RTL implementation of a complete transformer decoder layer, developed by Sigmantic AI using their autonomous verification agents. It runs Qwen2.5-0.5B at 0.56 tok/s on actual FPGA hardware, which is 140x faster than when first deployed. All blocks are bit-exact-verified against a NumPy golden model prior to being released.
Why a vibe-coder should care
It maintains constant inference throughput despite long conversations due to compressing the KV cache on the hardware. It’s completely open source and reproducible from a git clone, which is practically unheard of in AI hardware.
▌ More finds