GitHub radar

A local coding model that thinks less and answers faster

UkisAI retrained the open Qwen3.8-27B to think less before answering and work faster on code and agent tasks. A compressed version for home computers is available too.

01ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF 177231k downloads/mo27B paramstext-generation

This is a local version of the Swift 1.5 model, released by the startup UkisAI on top of the open Qwen3.8-27B from Alibaba. The authors retrained it to think less before answering: by their own numbers, it spends 58.5% fewer tokens on reasoning, while the final score is even slightly higher than the original Qwen. On some tasks this gives an almost 9.18-times speed gain. The model is tuned for coding and agent-style work, meaning multi-step tasks rather than plain chat. For local use the authors released a compressed version in several sizes, where a smaller file needs less memory at a small cost in accuracy. All of it is open, available without a subscription and without calling the vendor's own servers.

Why a vibe-coder should care

It lets you keep a coding and simple-agent model on your own computer, without a subscription and without a per-token bill. It will not replace a frontier model for serious work, but it fits playing with code, routine repeated tasks, or as a backup when you have no paid access at hand.

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up this model locally for me: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF — pick the compressed file that fits my memory, run it through llama.cpp or Ollama, and show me how to talk to it.

A regular computer with 16 GB of memory — the smallest version takes about 8-9 GB on disk, the more accurate one about 12 GB.

Open on Hugging Face