GitHub radar

Qwen's coding model shrunk 6x to run on your own machine

Researchers at ISTA-DASLab compressed Qwen's Qwen3.8-Flash-Next coding model almost six times by removing half its experts and shrinking the remaining weights. The result fits on a single 32 GB accelerator and keeps most of its quality on everyday coding tasks.

01ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF 172148k downloads/mo176.9B paramsimage-text-to-text

Qwen released Qwen3.8-Flash-Next, a 176.9-billion-parameter model built from many internal experts and meant for server use, weighing 354 GB in its original form. ISTA-DASLab, a lab at the Institute of Science and Technology Austria that works on compressing large models, released a code-focused version: half of the experts were removed entirely, and the remaining weights were recalculated at lower precision. The resulting GGUF file weighs 58.4 GB, roughly six times lighter than the original, and it fits in the memory of a single 32 GB accelerator. The authors state plainly how much quality was lost: on SWE-bench Verified, which tests long agentic work on real repositories, the model retains 91.3% of the base model's score. On LiveCodeBench, which covers shorter, simpler coding tasks, it retains 98.7%, so on single coding tasks the difference is barely noticeable. The model remains multimodal and still understands images, not just code, a capability the team tried to preserve through the cut.

Why a vibe-coder should care

You can run this one on your own hardware without a subscription or an internet connection. The quality is honest: on short coding tasks the gap with the full model is barely visible, while on long agentic work it falls behind by roughly a tenth. It is not a substitute for frontier work on something like Claude Code with Opus 5, but for simple, repeated coding tasks, or when there is no budget for a paid subscription, it is a workable option.

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up this coding model locally for me: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF - it's a ready GGUF file, run it through Ollama or llama.cpp (needs a Mac with 32 GB of memory or a 32 GB GPU), and show me how to give it coding tasks.

You need a Mac with at least 32 GB of unified memory, or a GPU with 32 GB of VRAM - that is how much space the compressed model takes up.

Open on Hugging Face