GitHub radar

NVIDIA Nemotron 3.5 Lightning: local agent via Ollama

NVIDIA's open 30B agent model with a 1M-token context window is now available locally via Ollama — designed to run continuously in the background.

01nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 360714k downloads/mo30B paramstext-generation

NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts model with only 3B parameters active at each inference step. It was built for always-on agents: personal assistants managing calendars and email, automation workflows, and similar long-running tasks. The context window is 1 million tokens. According to NVIDIA, the model delivers 4x higher throughput and 30% lower task completion time compared to similarly sized models. An MLX variant (23 GB) is available for Apple Silicon. Run locally with: ollama run nemotron-3.5-lightning.

Why a vibe-coder should care

An agent model with this context size and speed previously required server-grade GPUs. Via Ollama it runs on a Mac with 32 GB RAM or a 24 GB GPU — no cloud, no API key.

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Install NVIDIA Nemotron-3.5-Lightning locally: ollama run nemotron-3.5-lightning — and show me how to set it up as a background agent for long-running tasks

A Mac with 32 GB of RAM or more, or a GPU with 24 GB — the MLX version takes 23 GB.

Open on Hugging Face