As a founding member of the Open Secure AI Alliance for AI Safety and Security, Aible is a strong proponent of Open Models running fully under the control of the customer. While frontier models and open models both have their optimal use cases, customers often want fast, accurate models that they can easily post-train to specialize them to unique use cases.
We were really impressed by the performance of the base NVIDIA Nemotron 3.5 Lightning model in air-gapped environments. For long-running reasoning agents like AibleClaw, Nemotron 3.5 Lightning was over 57% faster than the comparably sized Qwen 3.5 35B a3b. Note that for this benchmark, Nemotron 3.5 Lightning was run on a single H100, while Qwen was run on two H100s. So this result was despite the Nemotron model running on half as many GPUs. Lightning also showed good scalability with an increasing number of agents running in parallel at the same time - delivering 50% higher per-user output speed.
Our most popular air-gapped solution is AibleClaw with NVIDIA Nemotron 3 Super model running on NVIDIA DGX Spark or the Dell Pro Max with GB10. Nemotron 3 Super is very capable on DGX Spark. But, we always wanted to post-train and run the smaller Nemotron 3 Nano 30B3A variant on DGX Spark so we could get even better performance. Post-training a smaller model on the output of a larger model can be a very effective way to get a faster, cheaper, specialized model for the task. On DGX Spark we potentially had the perfect solution - start with the Nemotron 3 Super and then use its output to post-train Nemotron 3 Nano model. But, as we previously reported, at that time, post-training sparse / Mixture of Expert (MoE) models was tougher than post-training dense models.
When we received early access to NVIDIA Nemotron 3.5 Lightning model, we were very eager to see how easy this model would be to post-train using the NVIDIA NeMo AutoModel based fine-tuning recipe that NVIDIA provided. We trained the model for just $0.80 (3.5 minutes of four NVIDIA A100 80GB GPUs) on only 400 rows of training data for a customer-specific text-to-sql use case and achieved a 20% improvement over the base model accuracy. When we had previously tried to post-train a similarly sized sparse / MoE model for $16.9 (148 minutes of a single NVIDIA H100 GPU) on 800 rows of the same dataset, the improvements were inconsequential. Thus, with the combination of NeMo AutoModel based fine-tuning recipe and the Nemotron 3.5 Lightning model, we were able to effectively post-train MoE models, with 21 times less cost with half as much data, and deliver significantly better accuracy improvement.
While these are our early test results with Nemotron 3.5 Lightning model, we are very encouraged by our initial tests of how easy it is to post-train. And of course being a smaller model, it runs even faster than Nemotron 3 Super on our NVIDIA DGX Spark setup. Post-training smaller open weights models is the obvious path to secure, accurate, performant, custom, sovereign AI.
Here is a live demonstration of an Aible agent using NVIDIA Nemotron 3.5 Lightning running completely locally on an NVIDIA DGX Spark desktop supercomputer. The video has not been edited or sped up in any way.

To see how NVIDIA Nemotron 3.5 Lightning performs on your own business use case with your data, request a briefing here.