The AI arena is free today

Open Superagent
NvidiaReleased on Jun 4, 2026

Nemotron 3 Ultra (550B A55B): Benchmarks, Pricing & Context Window

Nemotron 3 Ultra (550B A55B) is a language model from Nvidia, released in June 2026, with a 262K-token context window, and pricing from $0.500/M input, $0.100/M cached input, $2.20/M output.

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG. It uses a hybrid Latent Mixture-of-Experts (LatentMoE)

Input
TextImage
Output
Text

Nemotron 3 Ultra (550B A55B) benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Nemotron 3 Ultra (550B A55B) across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Nemotron 3 Ultra (550B A55B) holds up as conversations get longer.

Quality Tracker

Nemotron 3 Ultra (550B A55B) Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sun Sep 20 2026
Notice missing or incorrect data?

Nemotron 3 Ultra (550B A55B) pricing

Providers

Nemotron 3 Ultra (550B A55B) starts at $0.500 per million input tokens and $2.20 per million output tokens via DeepInfra. Reused prompt prefixes cost $0.100 per million cached input tokens.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.500$0.100$2.20262.1K/262.1K
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Nemotron 3 Ultra (550B A55B) model size

Nemotron 3 Ultra (550B A55B) has 550 billion parameters and was trained on 20 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
550BMoE
20Ttokens
36× tokens-to-params ratio
Frontier (200B+)
550B
1B7B70B405B

Nemotron 3 Ultra (550B A55B) context window

Input and output token limits for Nemotron 3 Ultra (550B A55B), plus how it ranks on long-context understanding.

InputOutput
262Ktokens
262Ktokens
394 pages of text
262K
8K128K1M

Try now

huggle
Nemotron 3 Ultra (550B A55B)in Huggle

Make it with
Nemotron 3 Ultra (550B A55B).

Nemotron 3 Ultra (550B A55B)

Nemotron 3 Ultra (550B A55B) latency

Nemotron 3 Ultra (550B A55B) time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Nemotron 3 Ultra (550B A55B) examples

Recent arena outputs from Nemotron 3 Ultra (550B A55B), picked from the highest-ranked matchups.

Nemotron 3 Ultra (550B A55B) license

Nemotron 3 Ultra (550B A55B) is released under the OpenMDW License v1.1 license, which permits commercial use, has 550.0B parameters, has a knowledge cutoff of September 2025.

License
OpenMDW License v1.1
Commercial use allowed
Parameters
550.0B
Knowledge cutoff
September 2025

Open Model, Data and Weights (OpenMDW) License Agreement, version 1.1 - permits commercial and non-commercial use.

Nemotron 3 Ultra (550B A55B) resources

Official sources for Nemotron 3 Ultra (550B A55B): provider documentation, paper or system card, source repository.

Nemotron 3 Ultra (550B A55B) vs other models

The most-compared alternatives to Nemotron 3 Ultra (550B A55B) are GPT-5 High, GPT-5.2 Pro, Claude Opus 4.5. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Nemotron 3 Ultra (550B A55B)

Models ranked just above and below Nemotron 3 Ultra (550B A55B) by LLM Stats score.

 

GPT-5 High

Score pending
 

GPT-5.2 Pro

Score pending
 

Claude Opus 4.5

Score pending
 

ERNIE 5.0

Score pending
 

Claude 3.7 Sonnet

Score pending
 

Grok Code Fast 1

Score pending

FAQ

Common questions about Nemotron 3 Ultra (550B A55B).

When was Nemotron 3 Ultra (550B A55B) released?

Nemotron 3 Ultra (550B A55B) was released on June 4, 2026 by Nvidia. This is the official Nemotron 3 Ultra (550B A55B) release date tracked on LLM Stats.

How much does Nemotron 3 Ultra (550B A55B) cost?

Nemotron 3 Ultra (550B A55B) pricing starts at $0.50 per million input tokens, $0.10 per million cached input tokens, $2.20 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Nemotron 3 Ultra (550B A55B)?

Nemotron 3 Ultra (550B A55B) has 550 billion parameters. It was trained on 20.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Nemotron 3 Ultra (550B A55B)?

Nemotron 3 Ultra (550B A55B) was created by Nvidia.

What is the license for Nemotron 3 Ultra (550B A55B)?

Nemotron 3 Ultra (550B A55B) is released under the OpenMDW License v1.1 license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Nemotron 3 Ultra (550B A55B)?

Nemotron 3 Ultra (550B A55B) has a knowledge cutoff of September 2025, meaning it was trained on data up to that point and may not know about events after it.

Where can I use Nemotron 3 Ultra (550B A55B)?

Nemotron 3 Ultra (550B A55B) is available through 1 provider including DeepInfra.

Where is the Nemotron 3 Ultra (550B A55B) paper or technical report?

Nemotron 3 Ultra (550B A55B) has a paper or technical report available at https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare Nemotron 3 Ultra (550B A55B) against?

Common Nemotron 3 Ultra (550B A55B) comparisons include Nemotron 3 Ultra (550B A55B) vs GPT-5 High, Nemotron 3 Ultra (550B A55B) vs GPT-5.2 Pro, Nemotron 3 Ultra (550B A55B) vs Claude Opus 4.5. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.