GitHub radar
DwarfStar: Local deepseek v4 flashing
Native inference engine for DeepSeek V4 Flash & GLM 5.2 by Salvatore Sanfilippo (Redis creator) — Apple Metal (Mac), NVIDIA CUDA (multi-gpu), AMD ROCm.
DwarfStar is a native inference engine developed by Salvatore Sanfilippo (Redis creator) for DeepSeek V4 Flash / DeepSeek V4 PRO / GLM 5.2. It is NOT a generic GGUF loader. It supports Apple Metal (Mac, main target), NVIDIA CUDA (multi-gpu) and AMD ROCm. It uses 2-bit quantization for routed MoE experts only (no shared experts, projections, routing), preserving quality while keeping the model size small. Built-in HTTP server allows micro-batching for multiple users. You can connect two MacBook M5 Max or M3 Ultra laptops with RDMA to do tensor parallelism. SSD streaming is possible if you don’t have enough RAM. See the benchmarks in README. For example, on MacBook Pro M5 Max 128GB it achieves 39 t/s. Beta.
Why a vibe-coder should care
If you have a Mac laptop with 96+ GB RAM (M5 Max or M3 Ultra), you can now run DeepSeek V4 Flash locally without any cloud API fees! Also works with older CUDA GPUs based on Ada Lovelace architecture (not supported anymore by vLLM for new models).
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Deploy DwarfStar from https://github.com/antirez/ds4 — follow the README, download the ds4f-q2 quantization with ./download_model.sh ds4f-q2, build with make, and show me how to run ./ds4. Ask if you need dependencies or permissions.
Requires a Mac with 96+ GB of RAM — for example, a MacBook Pro M5 Max or M3 Ultra. Machines with less RAM can use SSD streaming, but it's slower. The PRO model requires 512 GB.
Open on GitHub▌ More finds