GitHub radar

Introducing FreeToken — Access 290B+ MoE Models on a Gaming PC!

Edge-native Mixture of Expert inference server developed by Berkeley/MIT FlashML Team, allowing to deploy cutting-edge Open-Weight MoE models like DeepSeek-V4-Flash locally on consumer-grade NVIDIA RTX GPUs

01FlashML-org/FreeToken 6.8kPython

FreeToken is a device-native inference server that makes it possible to run large open-weight MoE (Mixture of Experts) models on edge devices. It was just released a month ago and has already got 3300 stars! Developed by Berkeley/MIT FlashML team (Matei Zaharia, co-author of Apache Spark, Ray, Song Han, Ion Stoica) and published on arXiv (2608.16157), it leverages bandwidth adaptive CPU–GPU co-execution, global LRU expert cache and semantic KV cache checkpoints to load models layer by layer to GPU and host memory, caching hot experts. With this approach, one can deploy 290B+ models such as DeepSeek-V4-Flash, Qwen3.6-35B-A3B or GLM-5.2 to NVIDIA RTX 30/40/50 gaming GPUs locally and avoid paying per token and uploading your data to the cloud. Supports Anthropic API out-of-the-box, so one can connect it to Claude Code and other tools like Codex and OpenCode. Desktop apps for Windows and Linux with UI are available, CLI (pip).

Why a vibe-coder should care

If you’ve got a gaming PC with NVIDIA RTX GPU, now you can run large open models locally and stop paying per token and uploading your data to the cloud. Connect FreeToken to Claude Code through Anthropic API

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Install FreeToken from here: https://github.com/FlashML-org/FreeToken — follow the README to install CLI (`pip install 'freetoken[accel]'`), download a model (e.g. Qwen3.6-35B-A3B) and start the OpenAI-compatible inference server on localhost. Ask me if you need any keys or permissions.

Requires an NVIDIA RTX 30, 40, or 50 series GPU — won't run on a Mac or a computer without a discrete GPU.

Open on GitHub