OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 3 MIN READ

NVIDIA Releases Nemotron 3.5 Lightning 30B MoE Model on Ollama for Local Agentic Work

NVIDIA's Nemotron 3.5 Lightning, a 30B-parameter Mixture-of-Experts model with 3B active parameters per token, is now available via Ollama for fully local inference. Optimized…

Stylized visualization of a Mixture-of-Experts model architecture running locally on hardware
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • The model uses a hybrid MoE architecture delivering up to 4x higher throughput than comparable open models through speculative decoding and multi-token prediction.
  • It targets practical agentic workloads including personal assistants, coding sub-agents and security operations that benefit from persistent local context without sending data externally.
  • Open weights trained on open datasets allow post-training for narrow tasks, with the same Ollama CLI and API supporting seamless switching between local and cloud models.

Model Architecture and Local Focus

According to the August 11 2026 Ollama blog post, Nemotron 3.5 Lightning is a 30 billion parameter model that activates only 3 billion parameters per token using a hybrid Mixture-of-Experts design. NVIDIA states this makes it suitable for local systems rather than datacenter-scale hardware.

The source highlights that the model runs on NVIDIA RTX PCs, RTX PRO workstations, DGX Spark, DGX Station, as well as in datacenters and cloud environments. For Apple silicon users, Ollama provides a nemotron-3.5-lightning:30b-mlx variant described as offering state-of-the-art performance on that platform.

Agentic Capabilities and Context

NVIDIA developed the model with the Nemotron Coalition specifically for agent harnesses, tool calling, instruction following, coding and multi-turn interactions. It supports a context length of up to 1 million tokens, which the source says accommodates long tool histories across extended workflows.

The announcement positions the model for tasks such as reading files, calling tools, sorting results and retrying failed steps. These operations often do not require the full capacity of larger models, making the sparse activation approach efficient for persistent agents.

Performance Claims and Optimizations

The source reports that Nemotron 3.5 Lightning provides 4x higher throughput and 30% faster task completion time compared to other leading open models of similar size. It attributes this to optimized inference using speculative decoding and multi-token prediction, with support for DFlash or DSpark.

NVIDIA claims leading accuracy on agentic, coding and reasoning tasks. Full benchmark results and test configurations are referenced to NVIDIA’s separate launch blog, which is not reproduced here. Throughput is emphasized as the key metric for long-running agents.

Use Cases and Customization

Intended applications listed include long-running personal assistants that handle email, calendar, projects and bookings while keeping all context local. Additional examples are coding sub-agents for testing and refactoring, and security operations such as alert enrichment, incident classification and log querying.

The model is described as open, trained on open datasets, enabling users to post-train it for specific narrow jobs and deploy the result from edge devices to datacenters. It can also function as a local tier alongside larger cloud models, routing only high-complexity steps externally while using the same CLI and API.

Getting Started with Ollama

The blog instructs users to download Ollama and run the model with commands such as ollama run nemotron-3.5-lightning for general chat. It lists several pre-configured agent launches including Claude Code, OpenClaw, Hermes Agent and OpenCode, all compatible with the same model.

The source notes that this pattern extends to Ollama’s cloud offerings, allowing agents to offload individual steps to larger models without altering the surrounding code.