GitHub radar

Kimi K3 in C: 2.78T-Param LLM on CPU using 8 GB of RAM

Runs Moonshot AI’s Kimi K3 (2.78T parameter mixture-of-experts language model) on a single CPU using only 8.24GB of RAM while streaming a 1.56TB checkpoint from disk. No GPUs, frameworks, just 176KB of C99 code.

C99 engine for running Moonshot AI’s Kimi K3 (2.78T parameter mixture-of-experts language model) fully on CPU without BLAS or CUDA and without an ML framework. Streams a 1.56TB checkpoint from NVMe and only maintains 8GB of active experts in memory. Gets identical results for all memory sizes between 8GB and 224GB.

Why a vibe-coder should care

Shows how modern AI can be run on commodity hardware (like your laptop) instead of specialty GPU clusters. Proves that with clever streaming and quantization techniques, trillion parameter inference is accessible to everyone’s home computer.

Open on GitHub