GitHub radar
A local coding model that thinks less and answers faster
UkisAI retrained the open Qwen3.8-27B to think less before answering and work faster on code and agent tasks. A compressed version for home computers is available too.
This is a local version of the Swift 1.5 model, released by the startup UkisAI on top of the open Qwen3.8-27B from Alibaba. The authors retrained it to think less before answering: by their own numbers, it spends 58.5% fewer tokens on reasoning, while the final score is even slightly higher than the original Qwen. On some tasks this gives an almost 9.18-times speed gain. The model is tuned for coding and agent-style work, meaning multi-step tasks rather than plain chat. For local use the authors released a compressed version in several sizes, where a smaller file needs less memory at a small cost in accuracy. All of it is open, available without a subscription and without calling the vendor's own servers.
Why a vibe-coder should care
It lets you keep a coding and simple-agent model on your own computer, without a subscription and without a per-token bill. It will not replace a frontier model for serious work, but it fits playing with code, routine repeated tasks, or as a backup when you have no paid access at hand.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Set up this model locally for me: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF — pick the compressed file that fits my memory, run it through llama.cpp or Ollama, and show me how to talk to it.
A regular computer with 16 GB of memory — the smallest version takes about 8-9 GB on disk, the more accurate one about 12 GB.
Open on Hugging Face▌ More finds