GitHub radar
Ling-3.0-tiny: Fast Reasoning on a MacBook
inclusionAI has announced the release of their new model Ling-3.0-tiny, which is a 7.9B MoE model with a thinking mode running at 86–90 tokens/sec on M4 Pro Macbook and using ~8GB RAM.
Ling-3.0-tiny is a hybrid MoE reasoning model developed by inclusionAI, which consists of 7.9B total parameters, but only activates 1.3B parameters per token. With its KDA-MLA architecture and 128-expert sparse MoE (8 routed + 1 shared expert active per token), it supports fast inference. It can run at 86–90 tokens per second on MacBook M4 Pro in FP8 mode, using ~8.34 GB RAM under an 8K context window. With YaRN scaling, the context window size can be scaled up to 256K tokens. When enable_thinking=true, temperature=1.0, top_p=0.95 is set, it will perform step-by-step reasoning first and then output the final answer. You can access this model through SGLang(Pre-built docker image), a special vLLM branch or an open pull request of Ollama for apple silicon.
Why a vibe-coder should care
This is a small but fast reasoning model, which can run on normal Macbook, perfect for situations where you want to think step by step but don’t have the money to pay for cloud computing resources. inclusionAI is a China based AI team. Please note that the support for Ollama is still in a pull request and not merged yet. At this moment, using SGLang with docker image is the most stable way to use this model.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Install Ling-3.0-tiny from inclusionAI for me: https://huggingface.co/inclusionAI/Ling-3.0-tiny — run it with SGLang (Docker) following the model card instructions and show me how to call it with thinking mode enabled (enable_thinking=true).
A regular laptop with 16 GB of RAM or more — in FP8 format the model takes up about 8 GB and delivers 86–90 tokens per second on a MacBook with an M4 Pro chip.
Open on Hugging Face▌ More finds