NVIDIA's own benchmarks show Nemotron 3.5 Lightning trailing Qwen 3.6 35B-A3B on 11 of 12 tests - yet it claims up to 4x faster token generation. That is the design: a 30B mixture-of-experts model with only 3B active parameters per token, built for the repetitive execution steps in an agent pipeline where speed and cost matter more than peak reasoning. The catch is that full-precision BF16 weights require an 80 GB GPU, and NVIDIA recommends the NVFP4 checkpoint for production use.
What Makes Nemotron 3.5 Lightning Different From Other 30B Models
Nemotron 3.5 Lightning uses a hybrid architecture that interleaves Mamba-2 state-space layers, MoE routing, and selective attention layers - NVIDIA tags this combination as nemotron_h. The Mamba-2 layers reduce the memory and compute overhead of long-context attention, which is how Lightning supports a 1M-token context window (practically capped at 256K on a single H100 80GB per the model card) without the quadratic cost a pure-attention model would face.
| Specification | Value |
|---|---|
| Total parameters | 30B |
| Active parameters per token | 3B |
| Architecture | Mamba-2 + MoE + Attention hybrid |
| Context length | Up to 1M tokens (256K on single H100) |
| Precision options | BF16, NVFP4 |
| License | OpenMDW v1.1 (commercially usable) |
| Languages | English, Spanish, French, German, Italian, Japanese + coding |
| Pre-training corpus | 20T+ tokens |
| Reasoning mode | Toggleable via enable_thinking in chat template |
| Recommended sampling | Temperature 1.0, top-p 0.95 |
NVIDIA positions Lightning as the smallest member of the Nemotron 3 family, trained specifically for agent harness behavior - tool calls, output validation, result formatting, and subagent delegation. A frontier reasoning model (like Nemotron 3 Ultra) handles complex planning, while Lightning handles routine execution steps that would otherwise burn through a premium model's token budget.
Benchmark Reality: Where Lightning Wins and Where It Loses
The HuggingFace model card publishes 14 benchmark rows comparing Lightning against Qwen 3.6 35B-A3B, Gemma 4 26B-A4B, Nemotron 3 Nano/Super, and GPT-OSS 20B. The 10 most decision-relevant rows are shown below:
| Benchmark | Nemotron 3.5 Lightning | Qwen 3.6 35B-A3B | Gemma 4 26B-A4B | GPT-OSS 20B |
|---|---|---|---|---|
| MMLU Pro | 81.94 | 85.63 | 85.20 | 76.40 |
| GPQA Diamond (no tools) | 75.44 | 83.40 | 79.61 | 71.46 |
| SWE-bench Verified | 51.56 | 70.12 | 57.40 | 52.44 |
| SWE-bench Multilingual | 39.33 | 63.40 | 43.40 | 41.93 |
| Terminal-Bench 2.1 | 24.58 | 44.38 | 37.22 | 15.17 |
| PinchBench | 85.37 | 88.07 | 74.70 | 57.20 |
| BrowseComp | 36.97 | 48.74 | 26.30 | - |
| IFBench (loose) | 71.88 | 63.71 | 77.25 | 68.50 |
| AA-LCR | 52.00 | 61.06 | 57.56 | 32.88 |
| SciCode | 32.60 | 35.33 | 40.28 | 38.63 |
Lightning trails Qwen 3.6 35B-A3B on 11 of the 12 comparable rows. The one exception is IFBench (instruction following, loose mode), where it scores 71.88 against Qwen's 63.71 - an 8-point lead that aligns with the model's agent-execution positioning: instruction adherence matters more than raw reasoning when handling tool-call formatting and output validation.
The speed story is where Lightning separates itself. On NVIDIA's PinchBench evaluation - 10,000 agent tasks measured in H100 GPU-hours - Lightning completes the workload in approximately 16.5 GPU-hours at 86% accuracy, versus 23.5-24 GPU-hours for Qwen 3.6 35B at 87% accuracy and 25-26 GPU-hours for Gemma 4 26B at 73% accuracy:
That is roughly 30% less compute for near-identical accuracy - significant if you are running 10,000 agent steps and paying per GPU-hour. NVIDIA also reports Lightning at approximately 670 output tokens/second with ~23-24 Intelligence Index points on the Artificial Analysis leaderboard, placing it on the Pareto frontier for open-weight models under 40B total parameters.
Community testing on r/LocalLLaMA produced similar speed findings: a DGX Spark user reported 78.5 tokens/second target-only and 90.7 tokens/second with speculative decoding on the NVFP4 checkpoint. Another user in the main discussion thread placed Lightning's quality "between Gemma 4 26B and 31B, much closer to 26B," noting it runs about 2x faster than Gemma 31B but generates many thinking tokens that offset the speed gain in practice. Early GGUF Q4 conversions reported 25 GB file sizes with unused-tensor warnings, suggesting the hybrid Mamba-2 architecture may not be fully supported in all GGUF-based tools yet.
Hardware Requirements and Local Deployment Options
The BF16 checkpoint - the full-precision reference weights - requires 1x H100 80GB or 1x A100 80GB for single-GPU deployment, with a 65.8 GB sharded safetensors file. NVIDIA explicitly recommends the separate NVFP4 release for production inference.
| Checkpoint | File size (approx.) | Minimum GPU | Best for |
|---|---|---|---|
| BF16 | 65.8 GB | 1x H100/A100 80GB | Fine-tuning, research, building quantized variants |
| NVFP4 | ~16 GB | RTX 5090, DGX Spark, Jetson (Blackwell/Hopper/Ampere) | Production inference, agent deployment |
| GGUF Q4 | ~25 GB | 16 GB+ VRAM (community-built) | llama.cpp, Ollama, LM Studio |
NVIDIA collaborated with four local-serving projects for Day-0 support: vLLM, SGLang, Ollama, and llama.cpp (with LM Studio and Unsloth also named). The model supports three speculative decoding strategies:
- DSpark - recommended for DGX Spark and low-concurrency data-center inference
- DFlash - alternative draft model, workload-dependent
- MTP (Multi-Token Prediction) - built into the model, best for medium-to-high concurrency
For hosted access without local hardware, Lightning is available through build.nvidia.com as a NIM microservice and on OpenRouter. The NVFP4 checkpoint runs on GeForce RTX 5090, DGX Spark, OEM GB10 systems, and NVIDIA Jetson.
The Agent Routing Pattern: Where Lightning Delivers Value
NVIDIA's open-source routing library, NeMo Switchyard, routes each agent workflow step to the best-fit model based on accuracy, speed, and cost. Plans route up to a frontier reasoning model; execution routes down to Lightning. NVIDIA's internal benchmark claims Switchyard maintained "frontier-level" task completion while reducing cost to approximately one-third of using Opus 4.8 alone. The execution tasks routed to Lightning include git pull, validating tool outputs, formatting results, and routine API calls.
This is backed by real usage. CodeRabbit published a real-world case on Reddit: they post-trained Nemotron 3.5 Lightning for code-review routing using targeted SFT plus RLVR (reinforcement learning with verifiable rewards), and it beat their baseline on a frozen 1,000-task evaluation - for under $100 in training costs. NVIDIA supports this workflow through NeMo Automodel and NeMo Megatron Bridge (LoRA/SFT), NeMo RL and NeMo Gym (reinforcement learning), and an open dataset called Nemotron-RL Agentic Terminal Pivot.
The practical question for teams building agent pipelines: can you fine-tune Lightning to handle the bulk of your agent steps at 3B active-parameter cost, reserving a frontier model for the few steps that need it?
Should You Use Nemotron 3.5 Lightning?
| Use case | Recommendation | Why |
|---|---|---|
| High-volume agent routing (tool calls, validation, formatting) | Lightning NVFP4 | 3B active params, 4x token speed, ~30% less compute at similar accuracy |
| General coding assistance | Qwen 3.6 35B-A3B | 70.12 vs 51.56 on SWE-bench Verified - 19-point gap |
| Complex reasoning / planning | Nemotron 3 Ultra or frontier model | Lightning is explicitly not the planning model |
| Single-GPU local chat on consumer hardware | Lightning NVFP4 on RTX 5090 | Works, but GGUF conversions are still rough; vLLM/SGLang more reliable |
| Fine-tuning for a narrow agent task | Lightning BF16 | 3B active params = cheaper to fine-tune; OpenMDW v1.1 permits commercial use |
| Instruction-following at scale | Lightning | IFBench 71.88 vs Qwen 63.71 - 8-point lead |
Lightning trades peak accuracy for speed in high-volume execution. The 30% compute reduction compounds over hundreds of agent steps per session, but the model launched August 11, 2026 and ecosystem maturity is still developing. The Mamba-2 hybrid architecture means GGUF-based tools may not fully support it yet. On NVIDIA hardware with vLLM or SGLang, the NVFP4 checkpoint is the safest deployment path today; on Apple Silicon or GGUF-only setups, wait for the community to stabilize conversions.
Frequently Asked Questions
Can Nemotron 3.5 Lightning run on consumer hardware?
Only the NVFP4 checkpoint is supported by NVIDIA for consumer hardware. The BF16 weights require an 80GB GPU (H100 or A100). NVFP4 runs on GeForce RTX 5090, DGX Spark, and NVIDIA Jetson. Community GGUF Q4 conversions exist for llama.cpp and Ollama on 16 GB+ VRAM systems, but early builds have reported unused-tensor issues with the Mamba-2 hybrid architecture.
How does Nemotron 3.5 Lightning compare with Qwen 3.6 35B?
Qwen 3.6 35B-A3B outperforms Lightning on 11 of 12 published benchmarks, with the largest gaps on SWE-bench Verified and Terminal-Bench. Lightning wins on IFBench instruction following and delivers approximately 30% faster task completion at similar accuracy on agent workloads. Qwen is the stronger general-purpose model; Lightning is the faster agent-execution model.
Is Nemotron 3.5 Lightning good for coding?
For raw coding benchmarks, no - SWE-bench Verified scores 51.56 versus Qwen 3.6's 70.12. However, CodeRabbit successfully fine-tuned Lightning for code-review routing at under $100, beating their baseline. The model is designed for customization: fine-tune it on your codebase conventions and it can handle routine code tasks at 3B active-parameter cost.
What license does Nemotron 3.5 Lightning use?
The model is released under OpenMDW v1.1 (Open Model Data Weight), which NVIDIA describes as released "as permissively as possible." The license permits commercial use, including weights, training data, and recipes. Full license terms are on the HuggingFace model card.