Nemotron 3.5 Lightning (free)
NVIDIA
nvidia/nemotron-3.5-lightning:freeNemotron 3.5 Lightning is NVIDIA's open 30B mixture-of-experts model with 3B active parameters and a 1M-token context window, built as a low-latency, low-cost execution layer for always-on agentic systems. This is the free-tier variant of the same model.
Cost rate
Free
Context
1M
Released
Aug 11, 2026
Input
Text
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
No top-ranked categories found.
Performance
Median latency and throughput measured across recent requests.
Throughput302 tok/s
Latency