GitHub radar
NVIDIA Nemotron 3.5 Lightning: local agent via Ollama
NVIDIA's open 30B agent model with a 1M-token context window is now available locally via Ollama — designed to run continuously in the background.
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts model with only 3B parameters active at each inference step. It was built for always-on agents: personal assistants managing calendars and email, automation workflows, and similar long-running tasks. The context window is 1 million tokens. According to NVIDIA, the model delivers 4x higher throughput and 30% lower task completion time compared to similarly sized models. An MLX variant (23 GB) is available for Apple Silicon. Run locally with: ollama run nemotron-3.5-lightning.
Why a vibe-coder should care
An agent model with this context size and speed previously required server-grade GPUs. Via Ollama it runs on a Mac with 32 GB RAM or a 24 GB GPU — no cloud, no API key.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Install NVIDIA Nemotron-3.5-Lightning locally: ollama run nemotron-3.5-lightning — and show me how to set it up as a background agent for long-running tasks
A Mac with 32 GB of RAM or more, or a GPU with 24 GB — the MLX version takes 23 GB.
Open on Hugging Face▌ More finds