rdwyer@fastmail.com · junk-mail
NVIDIA Nemotron 3.5 Lightning
Wed Aug 12, 2026 · 03:32 AM EDT
NVIDIA Nemotron 3.5 Lightning [www.nvidia.com/en-us/ai-data-science/f…nemotron/] is now available on
Ollama, and it runs completely on your own device. It’s a 30 billion parameter (3B active) open model from NVIDIA built for agents
that stay running: gathering context, calling tools, and working through multi-step tasks.
Nemotron 3.5 Lightning is made for agentic tasks such as reading a file, calling a tool, sorting a result, and retrying something
that failed. Most of these steps don’t need a large model. At 3B active parameters per token, it’s built for local systems rather
than the datacenter, and running locally means your data stays on your device.
Model highlights
* Runs where you work: 30B total parameters with only 3B active per token, on a hybrid Mixture-of-Experts architecture. It runs
locally on NVIDIA RTX PCs [www.nvidia.com/en-us/ai-on-rtx/], NVIDIA RTX PRO workstations
[www.nvidia.com/en-us/products/workstations/], NVIDIA DGX Spark
[www.nvidia.com/en-us/products/workstat…gx-spark/] and DGX Station
[www.nvidia.com/en-us/products/workstat…-station/], and in the datacenter and cloud.
* Built for agent harnesses: developed with the Nemotron Coalition and trained for the tools developers already use, across
coding, tool calling, instruction following and multi-turn work.
* 1M token context
* Optimized inference: speculative decoding using multi-token prediction (MTP), DFlash or DSpark, offering up to 4x higher
throughput than comparable open models.
* Yours to customize: an open model trained on open datasets. Post-train it for a specific task and run the result anywhere, from
edge to datacenter.
What you can build
Nemotron 3.5 Lightning excels on the following workloads:
* Long-running personal assistants. Email, calendar, projects and bookings. Running locally, the agent can use local context and
none of it is sent elsewhere.
* Coding sub-agents. Running tests, searching the codebase and applying refactors, inside the harnesses you already use.
* Security operations. Enriching alerts, classifying incidents, querying logs, correlating indicators and preparing structured
findings for analysts.
* A local tier alongside the cloud. Nemotron 3.5 Lightning handles the high-volume steps locally, and a larger hosted model picks
up the few that need one. Same CLI, same API.
* A specialist you train yourself. Open weights and open datasets, so you can post-train Nemotron 3.5 Lightning for one narrow
job and run the result locally.
Get started
Download Ollama [ollama.com/download], then run Nemotron 3.5 Lightning with your tool of choice.
General chat
ollama run nemotron-3.5-lightning
Claude Code
ollama launch claude --model nemotron-3.5-lightning
OpenClaw
ollama launch openclaw --model nemotron-3.5-lightning
Hermes Agent
ollama launch hermes --model nemotron-3.5-lightning
OpenCode
ollama launch opencode --model nemotron-3.5-lightning
For users on Apple silicon, Ollama offers the model with state-of-the-art performance: nemotron-3.5-lightning:30b-mlx.
See more integrations [ollama.com/library/nemotron-3.5-lightning] on the model page.
The same pattern works for models running in Ollama’s cloud, so an agent can send an individual step to a larger model without
changing anything else.
If you have any feedback, reply directly to this email or join Ollama's Discord [discord.gg/ollama].
❤️ Ollama
You are receiving this email because you opted in to receive updates from Ollama
Ollama, 744 High Street, Palo Alto, CA 94301
Unsubscribe [app.loops.so/unsubscribe/cmspru0jhiuv1…43b493db4]