NVIDIA Expands Agentic AI Portfolio With Nemotron 3.5 Lightning and NeMo Switchyard

NVIDIA Expands Agentic AI Portfolio With Nemotron 3.5 Lightning and NeMo Switchyard

New open model and intelligent routing technology target faster, more cost-efficient AI agents across PCs, edge devices, data centres and cloud environments

NVIDIA has expanded its agentic AI portfolio with the introduction of Nemotron 3.5 Lightning, an open model designed for high-volume AI agent workloads, alongside NeMo Switchyard, an open-source model-routing library aimed at helping enterprises improve the performance, cost and efficiency of increasingly complex AI systems.

The releases come as enterprise artificial intelligence moves beyond standalone chatbots toward autonomous and always-on agents capable of planning tasks, using tools and coordinating multiple models across extended workflows.

Rather than relying on a single large AI model for every request, this emerging architecture uses a system of models, with different models selected according to the complexity and requirements of each task.

NVIDIA is positioning Nemotron 3.5 Lightning and NeMo Switchyard as complementary technologies for this environment: one provides a smaller, specialized model for high-volume workloads, while the other determines which model should handle each request.

Nemotron 3.5 Lightning Targets High-Volume AI Workloads

Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model developed for specialized tasks within larger multi-agent systems.

According to NVIDIA, the model can deliver up to four times faster output speeds, resulting in 30% faster agentic task completion compared with other models in its class.

The model is intended for tasks including code review, tool use, security-alert monitoring and customer-service functions, where organizations may need to process large volumes of requests without relying on a larger frontier model for every interaction.

NVIDIA says the model combines high levels of reasoning capability with a smaller and more customizable architecture, allowing enterprises to adapt it to specific industries, datasets and workflows.

Organizations can post-train Nemotron 3.5 Lightning using NVIDIA NeMo and their own domain data, tools and operational processes.

The model was developed with contributions from the Nemotron Coalition, whose members provided datasets, inference software and evaluation methodologies.

From One AI Model to a System of Models

The launch reflects a broader change in the architecture of enterprise AI.

As AI agents become capable of carrying out longer and more complex workflows, organizations are increasingly combining different models rather than sending every request to the same system.

A powerful reasoning model, for example, may handle planning and complex decision-making, while smaller specialized models execute repetitive or domain-specific tasks.

This approach can potentially reduce computing requirements and operating costs while allowing enterprises to select models according to accuracy, speed, privacy and deployment requirements.

Nemotron 3.5 Lightning is designed to serve as one of these specialized models.

It can operate across local and enterprise infrastructure, including NVIDIA RTX PCs, DGX Spark, DGX Station and Jetson systems, as well as RTX PRO workstations, data centres, edge environments and cloud infrastructure.

Local and on-premises deployment can also provide organizations with greater control over sensitive data and AI workloads.

NeMo Switchyard Introduces Intelligent Model Routing

Alongside the new model, NVIDIA introduced NeMo Switchyard, an open-source routing library designed to automatically select the most appropriate AI model for individual tasks within an agent workflow.

The technology addresses a growing challenge for organizations deploying multiple AI models.

Some models perform better at coding, others at reasoning, while smaller models may be sufficient for routine requests. Using a powerful model for every task can increase costs, while manually determining which model should process each request adds complexity to AI applications.

NeMo Switchyard is designed to automate that decision.

Developers can configure its routing mechanisms according to priorities such as quality, latency and cost, while maintaining their existing combination of open-source, proprietary and NVIDIA models.

Importantly, NVIDIA says developers can introduce Switchyard without having to rewrite their applications.

According to NVIDIA’s internal benchmarking, Switchyard maintained performance approaching frontier-model levels while reducing task-completion costs to approximately one-third of using a single high-end model for all requests.

As these figures are based on NVIDIA’s own testing, performance in production environments will depend on workloads, model combinations and deployment configurations.

Enterprises Test the Multi-Model Approach

A number of technology companies and enterprises are already evaluating or integrating NVIDIA’s model and routing technologies.

Cybersecurity company CrowdStrike is among organizations customizing Nemotron 3.5 Lightning for specialized workloads, while legal AI company Harvey, CodeRabbit and other AI developers are working with the model for domain-specific applications.

NVIDIA is also working with companies across the AI ecosystem to integrate NeMo Switchyard into existing development environments.

Among the reported results, LangChain said routing only a small proportion of calls to a frontier model reduced costs substantially in its multi-turn agent testing, although this came with an accuracy trade-off.

Ramp reported achieving comparable performance to a frontier model in its software-engineering benchmark while reducing both costs and runtime.

Other companies working with or evaluating Switchyard include Boomi, Cadence, Cognition, Kong, LiteLLM, Nous Research and Siemens.

These examples point toward a potentially important shift in enterprise AI economics: the most effective AI system may not necessarily be the one using the largest model for every request.

Greater Control Over AI Deployment

Another key element of NVIDIA’s strategy is enterprise control.

Nemotron 3.5 Lightning can be customized and deployed locally, on premises, at the edge, in data centres or through cloud infrastructure. This flexibility could be particularly relevant for organizations operating in regulated industries or handling sensitive corporate information.

NVIDIA is also releasing training information and datasets where licensing permits.

Alongside Nemotron 3.5 Lightning, the company is publishing the Nemotron-RL-Agentic-Terminal-Pivot reinforcement-learning dataset used in post-training for coding-agent capabilities.

The approach is intended to provide greater transparency around model development while allowing researchers and enterprises to build upon the technology.

Agentic AI Enters Its Efficiency Phase

The significance of NVIDIA’s latest releases extends beyond the introduction of another AI model.

The rapid development of generative AI has largely focused on increasing model capability. As businesses move toward deploying autonomous agents at scale, however, efficiency, routing, infrastructure costs and deployment flexibility are becoming equally important considerations.

An always-on enterprise agent could potentially execute thousands or millions of individual tasks. Sending every one of those tasks to the most powerful—and most expensive—available model is unlikely to be economically efficient.

Intelligent routing offers an alternative: use sophisticated models when their capabilities are required and smaller specialized models when they can perform the task effectively.

Nemotron 3.5 Lightning and NeMo Switchyard represent NVIDIA’s approach to that emerging architecture.

For enterprises, the next phase of agentic AI may therefore be defined not simply by which model is the most powerful, but by how effectively multiple models, tools and computing environments can work together as one intelligent system.

Share