GitHub radar

TurboFieldfare is a 26B Model on 8 GB MacBook

An optimized Swift + Metal inference runtime for Google’s open Gemma 4 26B-A4B model that uses just ~2GB of memory — including 8GB Macs!

01drumih/turbo-fieldfare 6.2kSwift

Optimized Swift + Metal inference runtime for Google’s open Gemma 4 26B-A4B model. It’s designed specifically for this model (not a generic engine like MLX or llama.cpp), and only caches the shared 1.35GB core and FP16 KV cache in memory while streaming in experts from SSD as needed. This enables running the full 14.3 GB model with just ~2 GB of memory, making it accessible even on low-end Apple Silicon Macs. A fully documented set of 103 experimental results are included for various kernels, caching, I/O, and decodes. Includes a Mac native app, CLI, model streamer, and a local OpenAI server. Requires macOS 26+.

Why a vibe-coder should care

Run a true local OpenAI-compatible model on your Mac without cloud, subscriptions, or any data egress — using the Mac you have right now!

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Build TurboFieldfare from https://github.com/drumih/turbo-fieldfare (swift build -c release), launch the app and let the model download on first run.

Any M-chip Mac — even 8 GB of RAM is enough; the model takes ~15 GB on disk. Requires macOS 26+.

Open on GitHub