- Organizations
- Nvidia
- Nemotron 3.5 Lightning (30B A3B)
Nemotron 3.5 Lightning (30B A3B): Benchmarks, Pricing & Context Window
Nemotron 3.5 Lightning (30B A3B) is a language model from Nvidia, released in August 2026, with a 262K-token context window, and pricing from $0.080/M input, $0.040/M cached input, $0.200/M output.
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with ~3B active parameters, built for high-throughput specialized task execution in long-running / always-on agents. It ships with speculative decoding (MTP,
Nemotron 3.5 Lightning (30B A3B) benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Nemotron 3.5 Lightning (30B A3B) across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Nemotron 3.5 Lightning (30B A3B) holds up as conversations get longer.
Quality Tracker
Nemotron 3.5 Lightning (30B A3B) Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Nemotron 3.5 Lightning (30B A3B) pricing
Providers
Nemotron 3.5 Lightning (30B A3B) starts at $0.0800 per million input tokens and $0.200 per million output tokens via DeepInfra. Reused prompt prefixes cost $0.0400 per million cached input tokens.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0800 | $0.0400 | $0.200 | 262.1K/262.1K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Nemotron 3.5 Lightning (30B A3B) model size
Nemotron 3.5 Lightning (30B A3B) has 30 billion parameters. See how it compares to other models in the same parameter range.
Nemotron 3.5 Lightning (30B A3B) context window
Input and output token limits for Nemotron 3.5 Lightning (30B A3B), plus how it ranks on long-context understanding.
Try now
Make it with
Nemotron 3.5 Lightning (30B A3B).
Nemotron 3.5 Lightning (30B A3B) latency
Nemotron 3.5 Lightning (30B A3B) time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Nemotron 3.5 Lightning (30B A3B) examples
Recent arena outputs from Nemotron 3.5 Lightning (30B A3B), picked from the highest-ranked matchups.
Nemotron 3.5 Lightning (30B A3B) license
Nemotron 3.5 Lightning (30B A3B) is released under the OpenMDW License v1.1 license, which permits commercial use, has 30.0B parameters.
- License
- OpenMDW License v1.1
- Commercial use allowed
- Parameters
- 30.0B
Open Model, Data and Weights (OpenMDW) License Agreement, version 1.1 - permits commercial and non-commercial use.
Nemotron 3.5 Lightning (30B A3B) resources
Official sources for Nemotron 3.5 Lightning (30B A3B): provider documentation, official launch post, model weights.
Nemotron 3.5 Lightning (30B A3B) vs other models
The most-compared alternatives to Nemotron 3.5 Lightning (30B A3B) are Kimi K2 0905, Llama 3.1 Nemotron Ultra 253B v1, Qwen3 VL 235B A22B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.
- Nemotron 3.5 Lightning (30B A3B)vsKimi K2 0905
- Nemotron 3.5 Lightning (30B A3B)vsLlama 3.1 Nemotron Ultra 253B v1
- Nemotron 3.5 Lightning (30B A3B)vsQwen3 VL 235B A22B Instruct
- Nemotron 3.5 Lightning (30B A3B)vsClaude 3.5 Sonnet
- Nemotron 3.5 Lightning (30B A3B)vsQwen3 VL 32B Thinking
- Nemotron 3.5 Lightning (30B A3B)vsSarvam-105B
Models like Nemotron 3.5 Lightning (30B A3B)
Models ranked just above and below Nemotron 3.5 Lightning (30B A3B) by LLM Stats score.
FAQ
Common questions about Nemotron 3.5 Lightning (30B A3B).