Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to optimise agentic AI workflows
The 30-billion-parameter model and intelligent routing library offer greater control over data privacy and infrastructure, with early adopters already tailoring the technology for specific industry workloads.
Nvidia has expanded its Nemotron model family with the release of Nemotron 3.5 Lightning, a 30-billion-parameter open model designed for high-volume, long-running agentic AI workloads. Simultaneously, the company introduced NeMo Switchyard, an open-source library for intelligent model routing. The release, published on 11 August 2026, reflects a shift in enterprise AI strategy towards autonomous agents that require precise control over deployment, privacy, and operational efficiency.
Nemotron 3.5 Lightning is a mixture-of-experts model built to handle specialized tasks within larger multi-agent systems. Developed with contributions from the Nemotron Coalition, which provided evaluation methodologies and datasets, the model delivers up to four times faster output speeds. This performance improvement results in a 30 per cent faster completion rate for agentic tasks compared with other models in its class. The model is fully customizable and can be post-trained using Nvidia NeMo on organisational domain data to enhance accuracy for specific use cases.
The technology is designed to run across a diverse range of environments, from local AI systems such as Nvidia RTX PCs, DGX Spark, and Jetson devices, to edge AI devices, workstations, data centres, and cloud environments. This flexibility allows enterprises to maximise existing infrastructure investments while maintaining control over sensitive data. Nvidia also released Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset used to post-train the model for coding agent capabilities.
Several organisations are already customising Nemotron 3.5 Lightning for industry-specific workloads. CrowdStrike is utilising the model for cybersecurity, while Harvey with Trajectory is applying it to legal services. CodeRabbit with Baseten is using it for code review, Lila Sciences is enhancing reasoning for life sciences, and Fastino Labs is deploying it for software development, finance, and healthcare applications.
NeMo Switchyard addresses the complexity of managing multiple models by automatically directing prompts to the most capable and efficient model for each step of an agent workflow. Internal benchmarks indicate that the library maintains frontier-level accuracy while reducing task completion costs to nearly one-third of using Opus 4.8 alone. The library allows developers to tune routing algorithms based on priorities such as quality, latency, and cost without rewriting their applications.
Nemotron 3.5 Lightning is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an Nvidia NIM microservice, as well as through Nvidia Cloud Partners. NeMo Switchyard is available on GitHub and will be integrated into partner platforms in the near future.

