See how a local LLM thinks: Qwen3 running on your Mac
A new open source tool allows you to locally run Qwen3 and visualize the thought process layer by layer. Want to understand how LLMs work? Now you can!
From time to time, we come across a repository that brings an abstract concept into reality. jlens-qwen36 is such a repository. It locally runs a 27 billion parameter model (Qwen3.6) on your mac and allows you to visually explore its thought process layer by layer.
I’ve seen a lot of articles explaining how LLMs work. But none made it as clear as exploring this grid. So here’s what this tool is and why you should spend your afternoon playing with it.
Source
WeZZard/jlens-qwen36
★ 302Python
What the tool does
jlens-qwen36 is a visual debugger for local LLMs. It uses Apple’s MLX framework to load Qwen3.6-27B (4-bit quantized) locally on your mac. This means you can run the entire thing locally without relying on the cloud, an api key, or anything else. It then visualizes the model’s internals as a clickable grid of layers and token indices.
When you click on a cell, you get to see the top-10 words the neural net is currently focused on at that depth and index. Early layers are unintelligible but as you go deeper, you’ll notice the focus starts narrowing down on the correct answer. Seeing this process unfold is pure magic.
It uses a "Jacobian lens," an interpretability method that came out of Anthropic's research, adapted here to Qwen3 on consumer hardware. You don't need to know the math to enjoy the view.
Is a local LLM better than ChatGPT?
TLDR: No, a 27B model running on your laptop isn’t going to be as good as cutting edge cloud models in terms of quality. But this isn’t the metric that matters. The advantage of running a local LLM is that it’s private, subscription-free, has no rate limiting, and as we will see, it can be opened up and explored. This last point is much more valuable for educational purposes than a few benchmarks.
What's the catch with running it locally?
Hardware. Running a 27B model, even 4 bit, requires a mac with lots of unified memory. Think a decently specced M-series, not the lowest end version. Once you have a machine that can handle the model, running it locally is practically free aside from the cost of electricity.
- 1Runs on Apple Silicon via MLX — Mac only, for now.
- 24-bit quantized, so it fits in memory but still needs a healthy amount of it.
- 3Fully offline — good for privacy, and good for tinkering without burning API credits.
Who this is really for
If you’re an AI developer or prompt engineer, being able to see how the model transitions from “somewhat relevant words” to “the right answer” provides invaluable intuition that you won’t get by looking at any benchmark results. You begin to develop an intuition for what makes a good prompt versus one that gets lost. This intuition will make you a better prompt engineer, even if subconsciously.
You can now visually explore a cutting edge model using hardware you already own. A year ago this statement seemed preposterous.
How to run it yourself
There are two entry points. If you simply want to check out what it does, you can do so by clicking here and exploring the read-only demo, which runs entirely in the browser, no installation required. https://jlens.wezzard.com If you want to run it yourself, you’ll need an apple silicon mac with ~24GB of ram available.
git clone https://github.com/WeZZard/jlens-qwen36.git cd jlens-qwen36 && uv sync # pre-fitted lens (3.3 GB, two parts) — download and reassemble gh release download v0.2-fulldepth --repo WeZZard/jlens-qwen36 \ --pattern '*.npz.part-*' --dir data/lens/ cat data/lens/*.npz.part-* > data/lens/lens.npz && rm data/lens/*.part-* uv run python -m uvicorn jlens_qwen.serve:app --host 127.0.0.1 --port 8765 # open http://127.0.0.1:8765/
The first time you load the tool, it will automatically download the model (~15GB) from HuggingFace. So wait a minute. After that, visit the site and observe the grid fill up line by line as the model generates tokens. Click any cell to fixate it’s top 10 output.
A live window into a 27-billion-parameter model running on your own machine — you literally watch which words it's reaching for at each layer.
With over 300 stars in a few days, I know I’m not alone in wanting to peek behind the curtain. If you have a suitable mac, you’ll have the most fun exploring this tool to learn about LLMs.
Source: github.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles