The AI arena is free today

Open Superagent
NvidiaReleased on Mar 11, 2026

Nemotron 3 Super (120B A12B): Benchmarks, Pricing & Context Window

Nemotron 3 Super (120B A12B) is a language model from Nvidia, released in March 2026, with a 262K-token context window, and pricing from $0.085/M input and $0.400/M output.

Nemotron 3 Super is a 120B total / 12B active parameter hybrid Mamba-Attention Mixture-of-Experts model optimized for agentic reasoning, coding, planning, tool calling, and long-context analysis. It introduces LatentMoE (projecting tokens

Input
Text
Output
Text

Nemotron 3 Super (120B A12B) benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Nemotron 3 Super (120B A12B) across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Nemotron 3 Super (120B A12B) holds up as conversations get longer.

Quality Tracker

Nemotron 3 Super (120B A12B) Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Fri Sep 11 2026
Notice missing or incorrect data?

Nemotron 3 Super (120B A12B) pricing

Providers

Nemotron 3 Super (120B A12B) starts at $0.0850 per million input tokens and $0.400 per million output tokens via DeepInfra.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0850$0.400262.1K/262.1K
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Nemotron 3 Super (120B A12B) model size

Nemotron 3 Super (120B A12B) has 120 billion parameters and was trained on 25 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
120BMoE
25Ttokens
208× tokens-to-params ratio
Very large (80–200B)
120B
1B7B70B405B

Nemotron 3 Super (120B A12B) context window

Input and output token limits for Nemotron 3 Super (120B A12B), plus how it ranks on long-context understanding.

InputOutput
262Ktokens
262Ktokens
394 pages of text
262K
8K128K1M

Try now

huggle
Nemotron 3 Super (120B A12B)in Huggle

Make it with
Nemotron 3 Super (120B A12B).

Nemotron 3 Super (120B A12B)

Nemotron 3 Super (120B A12B) latency

Nemotron 3 Super (120B A12B) time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Nemotron 3 Super (120B A12B) examples

Recent arena outputs from Nemotron 3 Super (120B A12B), picked from the highest-ranked matchups.

Nemotron 3 Super (120B A12B) license

Nemotron 3 Super (120B A12B) is released under the NVIDIA Open Model License Agreement license, which permits commercial use, has 120.0B parameters, has a knowledge cutoff of June 2025.

License
NVIDIA Open Model License Agreement
Commercial use allowed
Parameters
120.0B
Knowledge cutoff
June 2025

NVIDIA Open Model License Agreement

Nemotron 3 Super (120B A12B) resources

Official sources for Nemotron 3 Super (120B A12B): provider documentation, paper or system card, official launch post, source repository, model weights.

Nemotron 3 Super (120B A12B) vs other models

The most-compared alternatives to Nemotron 3 Super (120B A12B) are Grok-3 Mini, K-EXAONE-236B-A23B, LongCat-Flash-Thinking. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Nemotron 3 Super (120B A12B)

Models ranked just above and below Nemotron 3 Super (120B A12B) by LLM Stats score.

 

Grok-3 Mini

Score pending
 

K-EXAONE-236B-A23B

Score pending
 

LongCat-Flash-Thinking

Score pending
 

LongCat-Flash-Lite

Score pending
 

Command A+

Score pending
 

DeepSeek-V3.2 (Thinking)

Score pending

FAQ

Common questions about Nemotron 3 Super (120B A12B).

When was Nemotron 3 Super (120B A12B) released?

Nemotron 3 Super (120B A12B) was released on March 11, 2026 by Nvidia. This is the official Nemotron 3 Super (120B A12B) release date tracked on LLM Stats.

How much does Nemotron 3 Super (120B A12B) cost?

Nemotron 3 Super (120B A12B) pricing starts at $0.09 per million input tokens and $0.40 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Nemotron 3 Super (120B A12B)?

Nemotron 3 Super (120B A12B) has 120 billion parameters. It was trained on 25.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Nemotron 3 Super (120B A12B)?

Nemotron 3 Super (120B A12B) was created by Nvidia.

What is the license for Nemotron 3 Super (120B A12B)?

Nemotron 3 Super (120B A12B) is released under the NVIDIA Open Model License Agreement license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Nemotron 3 Super (120B A12B)?

Nemotron 3 Super (120B A12B) has a knowledge cutoff of June 2025, meaning it was trained on data up to that point and may not know about events after it.

Where can I use Nemotron 3 Super (120B A12B)?

Nemotron 3 Super (120B A12B) is available through 1 provider including DeepInfra.

Where is the Nemotron 3 Super (120B A12B) paper or technical report?

Nemotron 3 Super (120B A12B) has a paper or technical report available at https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Super-Technical-Report.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare Nemotron 3 Super (120B A12B) against?

Common Nemotron 3 Super (120B A12B) comparisons include Nemotron 3 Super (120B A12B) vs Grok-3 Mini, Nemotron 3 Super (120B A12B) vs K-EXAONE-236B-A23B, Nemotron 3 Super (120B A12B) vs LongCat-Flash-Thinking. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.