GitHub radar
Colibrì: Train Trillion Parameter AI Models on your own computer!
Colibrì is a pure-C inference engine that makes storage, RAM and GPU one level of memory, enabling you to deploy cutting edge MoE models with up to 2.8 trillion parameters on your personal computer.
Colibrì is a pure C inference engine (no engine dependencies, single file core!) for cutting edge MoE models on your personal computer. It merges storage, RAM and VRAM into one level of memory: the dense shared core (~17B params, ~9.9 GB at int4) remains in RAM, and 19,456 routed experts (~19 MB each at int4, ~370 GB on disk) are streamed as needed with a per layer LRU cache. Models supported: Kimi K3 (2.8T) by Moonshot AI, GLM-5.2 (744B), Inkling (975B), DeepSeek V4 Flash (284B), Qwen3.6 (35B-A3B), OLMoE (7B). Built-in OpenAI compatible server allows clients speaking the OpenAI API (e.g., Claude Code by Anthropic) to connect to a locally hosted model. Backends: CPU, CUDA, Metal, NUMA. Web dashboard featuring realtime stats and live visualization of the entire set of 19,456 experts.
Why a vibe-coder should care
Deploy models that used to require datacenter GPUs on your personal computer, saving on API fees and ensuring your data stays private. Guaranteed performance degradation when not enough RAM is available ensures that you always know what kind of quality you’re getting.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Install colibri from https://github.com/JustVugg/colibri (grab the prebuilt release for my OS), help me pick and download a model per docs/quickstart.md, and start ./coli web.
A regular laptop with 10 GB of RAM or more — an M-chip Mac works. No GPU required, but you'll need a large SSD: models take up hundreds of gigabytes.
Open on GitHub▌ More finds