GitHub radar

Qwen3.8-27B GGUF Local run vision

Qwen3.8-27B (Alibaba’s 27B multimodal model) converted to Dynamic 3.0 GGUF quantization. 12Q, 262K context, image understanding, out-of-the-box for Ollama on your laptop!

01unsloth/Qwen3.8-27B-GGUF 2.8k7009k downloads/mo27B paramsimage-text-to-text

We’ve made Qwen3.8-27B (Alibaba’s 27B multimodal model) available as Dynamic 3.0 GGUF quantizations for local use with Ollama and llama.cpp. This model has image understanding capabilities, natively supports a 262K context length (can be scaled up to 1M tokens with RoPE), and is tuned for agency and tool invocation. With Dynamic 3.0 you can select from 12 different quantization levels (from IQ1, ~6GB to BF16, ~55GB) to find the right tradeoff between performance and hardware requirements. By keeping key weights less compressed, we believe we’ve found a way to surpass other quantizations at similar sizes in terms of accuracy.

Why a vibe-coder should care

A convenient way to run Qwen3.8-27B locally. Download a single file, put it in Ollama, and have a vision model that can process images and documents right on your device, without ever needing to go online. Even the 2-bit version works great on a 16GB Mac without requiring a GPU!

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Install Qwen3.8-27B by Alibaba locally via Ollama: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF — pick a suitable 4-bit GGUF for my Mac, install it and show me how to send a first prompt with an image.

The 4-bit version takes around 16 GB — a Mac with 32 GB of RAM or more, or a GPU with 24 GB is needed. The 2-bit version (~7 GB) runs on a Mac with 16 GB or more with some quality loss.

Open on Hugging Face