GitHub radar
Ornith-1.5-35B: MoE Model for Code and Agents
Ornith-1.5-35B-A3B (Mixture of Experts) by Deep Reinforce (Ornith Team) is a model for coding and agency, achieving 79% on SWE-bench Verified (compared to 73.4% for Qwen3.6-35B and 52% for Gemma 4-31B), and 70.2% on the MCP-Atlas agentic benchmark (compared to 62.8% for Qwen3.6-35B).
Ornith-1.5-35B-A3B is a Mixture of Experts model trained by Deep Reinforce (Ornith Team) for coding and agency. It has 35 billion parameters, but only uses approximately 3 billion at each token, so it runs faster than an equivalently sized dense model. It achieves 79% on SWE-bench Verified, compared to 73.4% for Qwen3.6-35B and 52% for Gemma 4-31B, and 70.2% on the MCP-Atlas agentic benchmark, compared to 62.8% for Qwen3.6-35B. It has a context length of 262,144 tokens, which can be increased to ~1M using RoPE scaling. It can make function calls through XML compatible with the OpenAI API, and use <think> blocks for chain of thought reasoning. Licensed under the MIT License. There is a prebuilt GGUF version here: ornith-ai/Ornith-1.5-35B-A3B-GGUF
Why a vibe-coder should care
This model may be useful if you’re running a local coding assistant/agent and need a model that performs better than Qwen3.6-35B on both code and agency benchmarks. This model is licensed under the MIT License, so it can be used commercially. There is also a prebuilt GGUF version that makes it easy to deploy locally using Ollama.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Install Ornith-1.5-35B from Deep Reinforce locally: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF — pick the right quantized version for my Mac, run it with Ollama and show me how to use it.
A Mac with 32 GB of RAM or more, or a GPU with 24 GB — the 4-bit version takes around 22 GB.
Open on Hugging Face▌ More finds