NVIDIA’s Nemotron 3.5 Lightning, a small model for big AI workloads

NVIDIA released Nemotron 3.5 Lightning, a compact, text-only model specifically engineered for high-volume, low-latency execution in autonomous AI agent systems.

Nemotron 3.5 Lightning is a small open-weights model designed for the high-volume, fast execution tasks required by long-running AI agents. It is the smallest model in the Nemotron 3 family, yet it outperforms its weight class. According to Artificial Analysis, it delivers performance comparable to OpenAI’s gpt-oss-120b while having roughly one-quarter of its total parameters.

NVIDIA also released NeMo Switchyard, an intelligent routing system that acts as the traffic controller for agentic workflows. It sends the planning tasks to advanced frontier models and the execution tasks to Nemotron 3.5 Lightning. In one reported example, the company says its LangChain escalation-routing setup reduced costs by 74% by sending 93% of routine tasks to Nemotron 3.5 Lightning and using more expensive frontier models for complex cases. Actual savings will vary depending on the workload and model pricing.

Key features

  • Mixture-of-Experts (MoE) architecture
  • Total parameters: roughly 30 billion
  • Active parameters per token: roughly 3 billion
  • Context window: up to 1 million tokens. Standard deployments typically use a 256K context limit, while support for the full 1 million tokens requires an appropriately configured inference runtime and sufficient memory capacity.

Availability and hardware support

Nemotron 3.5 Lightning is released under the OpenMDW License Agreement (version 1.1), allowing users to use, modify, and distribute the model for both research and commercial applications.

It is available for deployment through the NVIDIA NIM catalog and Hugging Face. The model is optimized to run on NVIDIA Blackwell, Hopper (supporting NVFP4 and W4A16), and Ampere GPU generations.

Architecture: the Mamba-2 and MoE hybrid

While frontier models are typically reserved for high-level planning and complex reasoning, NVIDIA positions the 3.5 Lightning model as a sub-agent workhorse. Its primary function is to handle the repetitive, high-frequency tasks required in long-running agentic workflows, such as information retrieval, multi-document aggregation, and structured output generation.

The defining characteristic of Nemotron 3.5 Lightning is its hybrid architecture. Rather than relying solely on standard transformer blocks, the model uses an interleaved design consisting of Mamba-2, Mixture-of-Experts (MoE), and select Attention layers.

This hybrid approach tries to balance the linear scaling of state-space models (Mamba-2) with the high capacity of MoE. Out of its 30 billion parameters, only 3 billion are active during any single forward pass. This allows the model to maintain a high level of intelligence while delivering high inference speed typically associated with much smaller models.

To further refine its training signals, NVIDIA integrated Multi-Token Prediction (MTP) layers. These layers allow the model to predict multiple future tokens simultaneously, providing a richer training signal than standard next-token prediction models. To speed up generation during inference, they applied speculative decoding where a small model predicts multiple tokens at once and a larger model verifies them all in parallel. MTP improves the training objective, while speculative decoding reduces the number of expensive sequential forward passes needed during generation.

The Execution Layer concept

The true power of Nemotron 3.5 Lightning lies in its role within a multi-agent system, where different models can handle different parts of a workflow. Using an expensive frontier model for every task can increase both costs and latency, especially when many tasks are routine or repetitive. A more efficient approach is to use a capable frontier model, such as GPT-4o or Claude 3.5 Sonnet, as the Orchestrator to define goals, plan the workflow, and handle more complex decisions. Nemotron 3.5 Lightning can then serve as the Execution Layer, carrying out the repetitive tasks that do not require a frontier model.

This approach allows each model to perform the tasks it is best suited for, helping reduce costs and improve execution speed.

Training methodology and optimization

The development of Nemotron 3.5 Lightning followed a five-stage pipeline:

  1. Pre-training: The model was pre-trained on over 20 trillion tokens, using a mix of crawled and synthetic data covering code, math, science, and general knowledge.
  2. Multi-Token Prediction (MTP) training: A dedicated phase was used to align the Multi-Token Prediction layers with the base model’s distribution.
  3. Supervised Fine-Tuning (SFT): The model was refined on synthetic datasets focused on tool calling, structured outputs, and long-context retrieval.
  4. Reinforcement Learning (RL): NVIDIA used Group Relative Policy Optimization (GRPO) across multiple environments (math, code, science). This was executed using an asynchronous RL architecture that decouples training from inference, leveraging the MTP layers to accelerate rollout generation.
  5. Post-Training Quantization (PTQ): The final model underwent quantization using a variant of static MSE calibration (such as NVFP4) and W4A16) for routed and shared experts.

Evaluation

NVIDIA reports the following benchmark results comparing the 16-bit floating-point (BF16) and 4-bit quantized (NVFP4) checkpoints of Nemotron-3.5-Lightning:

Task   Nemotron-3.5-Lightning-30B-A3B-BF16Nemotron-3.5-Lightning-30B-A3B-NVFP4
MMLU Pro        81.9481.62
GPQA Diamond (no tools)  75.4475.57
SWE-bench Verified  51.5652.80
AA-LCR (Long Context)52.0049.19

Compressing the model to 4-bit precision (NVFP4) leads to a minor drop of 0.32 points on MMLU Pro (from 81.94 to 81.62), demonstrating that core reasoning and general knowledge remain unaffected. In contrast, a 2.81-point drop in long-context retrieval indicates that sequence tracking and attention mechanisms are more sensitive to lower numerical precision.

Nemotron 3.5 Lightning scores 13 on the Artificial Analysis Intelligence Index, which is a composite benchmark that measures performance across reasoning, knowledge, mathematics, and coding. The model is placed well above the median score of 8 among open-weight models of similar size.

The next three-panel comparative analysis evaluates the leading language models across three indices: Intelligence, Speed, and Cost per Task (see the next picture). Nemotron 3.5 Lightning is highlighted in bright green.

Intelligence, speed, and cost performance of Nemotron 3.5 Lightning (source: Artificial Analysis)

Frontier reasoning models lead the benchmark, with Claude Opus 5.5 scoring 58, followed by Claude Fable 5.1 and GPT-6 at 53, and GPT-6.1 Sol at 52. Although Nemotron 3.5 Lightning ranks lower in overall reasoning performance with a score of 13, it delivers the fastest processing speed and lowest cost per task.

This finding is backed by Artificial Analysis data published by NVIDIA, placing Nemotron 3.5 Lightning in the most attractive speed-versus-performance quadrant (see the image below).

Intelligence index and output speed (source: NVIDIA’s blog)

Conclusion

Nemotron 3.5 Lightning is an open, compact AI model designed specifically for the execution layer of long-running agents. It handles high-volume repetitive tasks such as tool calls, result validation, code review, and other specialized operations quickly and efficiently.

The new model marks a shift in the ML landscape from “bigger is better” to “smarter routing.” Rather than relying on a single large model for every step, developers can use model routing to match each task with the model best suited to handle it.

NVIDIA’s NeMo Switchyard is designed to support this approach by routing planning work to frontier models and execution work to Lightning when appropriate.

Read more

Other popular posts