The AI arena is free today

Open Superagent
NvidiaReleased on Aug 11, 2026

Nemotron 3.5 Lightning (30B A3B): Benchmarks, Pricing & Context Window

Nemotron 3.5 Lightning (30B A3B) is a language model from Nvidia, released in August 2026, with a 262K-token context window, and pricing from $0.080/M input, $0.040/M cached input, $0.200/M output.

NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with ~3B active parameters, built for high-throughput specialized task execution in long-running / always-on agents. It ships with speculative decoding (MTP,

Input
Text
Output
Text

Nemotron 3.5 Lightning (30B A3B) benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Nemotron 3.5 Lightning (30B A3B) across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Nemotron 3.5 Lightning (30B A3B) holds up as conversations get longer.

Quality Tracker

Nemotron 3.5 Lightning (30B A3B) Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Sep 09 2026
Notice missing or incorrect data?

Nemotron 3.5 Lightning (30B A3B) pricing

Providers

Nemotron 3.5 Lightning (30B A3B) starts at $0.0800 per million input tokens and $0.200 per million output tokens via DeepInfra. Reused prompt prefixes cost $0.0400 per million cached input tokens.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0800$0.0400$0.200262.1K/262.1K
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Nemotron 3.5 Lightning (30B A3B) model size

Nemotron 3.5 Lightning (30B A3B) has 30 billion parameters. See how it compares to other models in the same parameter range.

Parameters
30BMoE
Large (30–80B)
30B
1B7B70B405B

Nemotron 3.5 Lightning (30B A3B) context window

Input and output token limits for Nemotron 3.5 Lightning (30B A3B), plus how it ranks on long-context understanding.

InputOutput
262Ktokens
262Ktokens
394 pages of text
262K
8K128K1M

Try now

huggle
Nemotron 3.5 Lightning (30B A3B)in Huggle

Make it with
Nemotron 3.5 Lightning (30B A3B).

Nemotron 3.5 Lightning (30B A3B)

Nemotron 3.5 Lightning (30B A3B) latency

Nemotron 3.5 Lightning (30B A3B) time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Nemotron 3.5 Lightning (30B A3B) examples

Recent arena outputs from Nemotron 3.5 Lightning (30B A3B), picked from the highest-ranked matchups.

Nemotron 3.5 Lightning (30B A3B) license

Nemotron 3.5 Lightning (30B A3B) is released under the OpenMDW License v1.1 license, which permits commercial use, has 30.0B parameters.

License
OpenMDW License v1.1
Commercial use allowed
Parameters
30.0B

Open Model, Data and Weights (OpenMDW) License Agreement, version 1.1 - permits commercial and non-commercial use.

Nemotron 3.5 Lightning (30B A3B) resources

Official sources for Nemotron 3.5 Lightning (30B A3B): provider documentation, official launch post, model weights.

Nemotron 3.5 Lightning (30B A3B) vs other models

The most-compared alternatives to Nemotron 3.5 Lightning (30B A3B) are Kimi K2 0905, Llama 3.1 Nemotron Ultra 253B v1, Qwen3 VL 235B A22B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Nemotron 3.5 Lightning (30B A3B)

Models ranked just above and below Nemotron 3.5 Lightning (30B A3B) by LLM Stats score.

 

Kimi K2 0905

Score pending
 

Llama 3.1 Nemotron Ultra 253B v1

Score pending
 

Qwen3 VL 235B A22B Instruct

Score pending
 

Claude 3.5 Sonnet

Score pending
 

Qwen3 VL 32B Thinking

Score pending
 

Sarvam-105B

Score pending

FAQ

Common questions about Nemotron 3.5 Lightning (30B A3B).

When was Nemotron 3.5 Lightning (30B A3B) released?

Nemotron 3.5 Lightning (30B A3B) was released on August 11, 2026 by Nvidia. This is the official Nemotron 3.5 Lightning (30B A3B) release date tracked on LLM Stats.

How much does Nemotron 3.5 Lightning (30B A3B) cost?

Nemotron 3.5 Lightning (30B A3B) pricing starts at $0.08 per million input tokens, $0.04 per million cached input tokens, $0.20 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Nemotron 3.5 Lightning (30B A3B)?

Nemotron 3.5 Lightning (30B A3B) has 30 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Nemotron 3.5 Lightning (30B A3B)?

Nemotron 3.5 Lightning (30B A3B) was created by Nvidia.

What is the license for Nemotron 3.5 Lightning (30B A3B)?

Nemotron 3.5 Lightning (30B A3B) is released under the OpenMDW License v1.1 license. This is an open-source / open-weight license that permits self-hosting.

Where can I use Nemotron 3.5 Lightning (30B A3B)?

Nemotron 3.5 Lightning (30B A3B) is available through 1 provider including DeepInfra.

Where is the Nemotron 3.5 Lightning (30B A3B) paper or technical report?

Nemotron 3.5 Lightning (30B A3B) has a paper or technical report available at https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/. Use that source for architecture, training, release and evaluation details.

What models should I compare Nemotron 3.5 Lightning (30B A3B) against?

Common Nemotron 3.5 Lightning (30B A3B) comparisons include Nemotron 3.5 Lightning (30B A3B) vs Kimi K2 0905, Nemotron 3.5 Lightning (30B A3B) vs Llama 3.1 Nemotron Ultra 253B v1, Nemotron 3.5 Lightning (30B A3B) vs Qwen3 VL 235B A22B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.