GitHub radar

Ornith 1.5 35B GGUF: Run it on your computer.

The GGUF repo has quantized weights for running this model locally using Ollama or llama.cpp

01ornith-ai/Ornith-1.5-35B-A3B-GGUF 3181469k downloads/mo35B paramstext-generation

Ornith 1.5 35B A3B is an ensemble of experts model for code development by ornith-ai. Although it has 35B parameters, it only uses 3B parameters at inference time and so is actually much faster than a fully dense 35B model. Ornith 1.5 35B A3B achieves 79% on SWE-bench Verified, which is the first time a local model has surpassed this score. With its 262,144 token context length, Ornith 1.5 35B A3B can handle all your code at once in a single session. Out of the box, it supports tool invocations and integrates seamlessly into a code developing agency.

Why a vibe-coder should care

Previously, Ornith 1.5 35B was only available in a server-based format, requiring a data center to run. Now, thanks to the GGUF format, you can run Ornith 1.5 35B on your powerful Mac or consumer GPU. If 9B was not enough for you, here comes 35B!

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Install Ornith 1.5 35B locally: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF — pick the right quantized version for my machine, run it via Ollama, and show me how to give it a coding task

A Mac with 32 GB of RAM or more, or a GPU with 24 GB — the 4-bit version takes around 21 GB.

Open on Hugging Face