Sync a voice

rdwyer@fastmail.com · junk-mail

NVIDIA Nemotron 3.5 Lightning

Wed Aug 12, 2026 · 03:32 AM EDT

From
Ollama <hello@ollama.com>
To
rdwyer@fastmail.com

NVIDIA Nemotron 3.5 Lightning [www.nvidia.com/en-us/ai-data-science/f…nemotron/] is now available on
Ollama, and it runs completely on your own device. It’s a 30 billion parameter (3B active) open model from NVIDIA built for agents
that stay running: gathering context, calling tools, and working through multi-step tasks.

Nemotron 3.5 Lightning is made for agentic tasks such as reading a file, calling a tool, sorting a result, and retrying something
that failed. Most of these steps don’t need a large model. At 3B active parameters per token, it’s built for local systems rather
than the datacenter, and running locally means your data stays on your device.

Model highlights

* Runs where you work: 30B total parameters with only 3B active per token, on a hybrid Mixture-of-Experts architecture. It runs
locally on NVIDIA RTX PCs [www.nvidia.com/en-us/ai-on-rtx/], NVIDIA RTX PRO workstations
[www.nvidia.com/en-us/products/workstations/], NVIDIA DGX Spark
[www.nvidia.com/en-us/products/workstat…gx-spark/] and DGX Station
[www.nvidia.com/en-us/products/workstat…-station/], and in the datacenter and cloud.

* Built for agent harnesses: developed with the Nemotron Coalition and trained for the tools developers already use, across
coding, tool calling, instruction following and multi-turn work.

* 1M token context

* Optimized inference: speculative decoding using multi-token prediction (MTP), DFlash or DSpark, offering up to 4x higher
throughput than comparable open models.

* Yours to customize: an open model trained on open datasets. Post-train it for a specific task and run the result anywhere, from
edge to datacenter.

What you can build

Nemotron 3.5 Lightning excels on the following workloads:
* Long-running personal assistants. Email, calendar, projects and bookings. Running locally, the agent can use local context and
none of it is sent elsewhere.

* Coding sub-agents. Running tests, searching the codebase and applying refactors, inside the harnesses you already use.

* Security operations. Enriching alerts, classifying incidents, querying logs, correlating indicators and preparing structured
findings for analysts.

* A local tier alongside the cloud. Nemotron 3.5 Lightning handles the high-volume steps locally, and a larger hosted model picks
up the few that need one. Same CLI, same API.

* A specialist you train yourself. Open weights and open datasets, so you can post-train Nemotron 3.5 Lightning for one narrow
job and run the result locally.

Get started

Download Ollama [ollama.com/download], then run Nemotron 3.5 Lightning with your tool of choice.

General chat
ollama run nemotron-3.5-lightning

Claude Code
ollama launch claude --model nemotron-3.5-lightning

OpenClaw
ollama launch openclaw --model nemotron-3.5-lightning

Hermes Agent
ollama launch hermes --model nemotron-3.5-lightning

OpenCode
ollama launch opencode --model nemotron-3.5-lightning

For users on Apple silicon, Ollama offers the model with state-of-the-art performance: nemotron-3.5-lightning:30b-mlx.

See more integrations [ollama.com/library/nemotron-3.5-lightning] on the model page.

The same pattern works for models running in Ollama’s cloud, so an agent can send an individual step to a larger model without
changing anything else.

If you have any feedback, reply directly to this email or join Ollama's Discord [discord.gg/ollama].

❤️ Ollama

You are receiving this email because you opted in to receive updates from Ollama
Ollama, 744 High Street, Palo Alto, CA 94301
Unsubscribe [app.loops.so/unsubscribe/cmspru0jhiuv1…43b493db4]