YC Startup Magnitude Speeds Up Local Models For AI Agents
A YC startup has released Magnitude, an open-source engine that tunes local AI models to specific hardware and plugs into Claude Code and other agents for free.
A startup from YC S25 has released Magnitude, an open-source engine for running AI models locally. It tunes computation to the user's specific hardware instead of using the same code for every device. By the author's own numbers, on Apple Metal this gives a 92% boost in generation speed and a 9% boost in prompt processing compared with llama.cpp. On NVIDIA CUDA the gains are 19% and 23% respectively, and memory use per agent drops by 27%. The tool is free, licensed under Apache 2.0, and connects in one click to Claude Code, Codex, OpenCode, and other agents. For everything else, there is an OpenAI-compatible API.
On its own, this tool does not make open models smarter. It just speeds up what an open model can already do on a given piece of hardware, and for serious agentic work that usually is not enough. In my experience, free open models still cannot handle real agent work on par with the top paid ones. But for experiments, simple repetitive tasks, or situations where a subscription just is not in the budget, a faster, more memory-efficient local engine is genuinely useful. It does not pretend to replace an expensive model, which I appreciate. The one-click connection to Claude Code and Codex is interesting in its own right. You can quickly test what a local model can pull off on your own tasks without touching your main agent pipeline.
Source: github.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles