Thinking Machines' Inkling Small: Fewer Parameters, Stronger Benchmarks
Thinking Machines — the lab of former OpenAI CTO Mira Murati — has released Inkling Small: 276 billion parameters, 12 billion active, Apache 2.0.
Thinking Machines — the lab of former OpenAI CTO Mira Murati — has released Inkling Small. The model has 276 billion parameters, with only 12 billion active, yet it outperforms its larger predecessor on key benchmarks: GPQA Diamond — 89% vs. 87%, Humanity's Last Exam — 32% vs. 30%.
The model uses an average of 24,000 tokens per task — half as many as DeepSeek Flash (45,000). Weights are available on Hugging Face under the Apache 2.0 license, and fine-tuning is available through the Tinker Playground platform directly in the browser.
The model once again proves that the race for scale is not the only path to smarter AI. An efficient architecture can match the giants at a fraction of the compute cost — which changes the economics of development for everyone paying per token.
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles