The AI arena is free today

Open Superagent

Llama 3.1 Nemotron Ultra 253B v1 vs Llama 4 Maverick

Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick are closely matched at 19.1 and 14.7 on the LLM Stats Score.

NVIDIA · Meta · Updated for 2026

Which is better?

Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick are closely matched on the overall LLM Stats Score at 19.1 and 14.7.

In the 2 individual benchmarks reported for both models, Llama 3.1 Nemotron Ultra 253B v1 wins 2; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose Llama 3.1 Nemotron Ultra 253B v1

  • you value its reported benchmark strengths — it wins 2 of 2 exact shared results
  • you want the most recent training data — it shipped Apr 2025

Choose Llama 4 Maverick

  • you want predictable pricing at $0.17/M input and $0.60/M output

At a glance

The differences that matter most.

Core performance indexes
19.1
#205
14.7
#233
18.8
#204
14.3
#229
12.4
#144
3.0
#218
Cost, coverage & limits
Benchmark wins
2 of 2
0 of 2
Input price
— / M
$0.17 / M
Output price
— / M
$0.60 / M
Context window
1,000,000

Capability indexes

Additional strengths measured across groups of related public benchmarks

1 shared
Index
Llama 3.1 Nemotron Ultra 253B v1
Llama 4 Maverick
15.6#214
20.4#157
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

6 reported for Llama 3.1 Nemotron Ultra 253B v1 · 13 for Llama 4 Maverick

2 shared

Llama 3.1 Nemotron Ultra 253B v1 outperforms in 2 benchmarks (GPQA, LiveCodeBench), while Llama 4 Maverick is better at 0 benchmarks.

Llama 3.1 Nemotron Ultra 253B v1 significantly outperforms across most benchmarks.

Mon Sep 14 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Model Size

Parameter count comparison

147.0B diff

Llama 4 Maverick has 147.0B more parameters than Llama 3.1 Nemotron Ultra 253B v1, making it 58.1% larger.

NVIDIA
Llama 3.1 Nemotron Ultra 253B v1
253.0Bparameters
Meta
Llama 4 Maverick
400.0Bparameters
253.0B
Llama 3.1 Nemotron Ultra 253B v1
400.0B
Llama 4 Maverick

Context Window

Maximum input and output token capacity

Only Llama 4 Maverick specifies input context (1,000,000 tokens). Only Llama 4 Maverick specifies output context (1,000,000 tokens).

NVIDIA
Llama 3.1 Nemotron Ultra 253B v1
Input- tokens
Output- tokens
Meta
Llama 4 Maverick
Input1,000,000 tokens
Output1,000,000 tokens
Mon Sep 14 2026 • llm-stats.com

Input capabilities

Documented input modalities across available providers

Llama 4 Maverick supports multimodal inputs, whereas Llama 3.1 Nemotron Ultra 253B v1 does not.

Llama 4 Maverick can handle both text and other forms of data like images, making it suitable for multimodal applications.

Llama 3.1 Nemotron Ultra 253B v1

Text
Images
Audio
Video

Llama 4 Maverick

Text
Images
Audio
Video

License

Usage and distribution terms

Llama 3.1 Nemotron Ultra 253B v1 is licensed under Llama 3.1 Community License, while Llama 4 Maverick uses Llama 4 Community License Agreement.

License differences may affect how you can use these models in commercial or open-source projects.

Llama 3.1 Nemotron Ultra 253B v1

Llama 3.1 Community License

Open weights

Llama 4 Maverick

Llama 4 Community License Agreement

Open weights

Release Timeline

When each model was launched

Llama 3.1 Nemotron Ultra 253B v1 was released on 2025-04-07, while Llama 4 Maverick was released on 2025-04-05.

Llama 3.1 Nemotron Ultra 253B v1 is 0 month newer than Llama 4 Maverick.

Llama 3.1 Nemotron Ultra 253B v1

Apr 7, 2025

1.4 years ago

2d newer
Llama 4 Maverick

Apr 5, 2025

1.4 years ago

Knowledge Cutoff

When training data ends

Llama 3.1 Nemotron Ultra 253B v1 has a documented knowledge cutoff of 2023-12-01, while Llama 4 Maverick's cutoff date is not specified.

We can confirm Llama 3.1 Nemotron Ultra 253B v1's training data extends to 2023-12-01, but cannot make a direct comparison without Llama 4 Maverick's cutoff date.

Llama 3.1 Nemotron Ultra 253B v1

Dec 2023

Llama 4 Maverick

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick side-by-side, then vote on the output you prefer.

Llama 3.1 Nemotron Ultra 253B v1
✓ Preferred
Llama 4 Maverick
Open in Playground

FAQ

Common questions about Llama 3.1 Nemotron Ultra 253B v1 vs Llama 4 Maverick.

Which is better, Llama 3.1 Nemotron Ultra 253B v1 or Llama 4 Maverick?

Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick are closely matched on the LLM Stats Score at 19.1 and 14.7. Llama 3.1 Nemotron Ultra 253B v1 is made by NVIDIA and Llama 4 Maverick is made by Meta. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does Llama 3.1 Nemotron Ultra 253B v1 compare to Llama 4 Maverick in benchmarks?

Llama 3.1 Nemotron Ultra 253B v1 scores MATH-500: 97.0%, IFEval: 89.5%, GPQA: 76.0%, BFCL v2: 74.1%, AIME 2025: 72.5%. Llama 4 Maverick scores DocVQA: 94.4%, MGSM: 92.3%, ChartQA: 90.0%, MMLU: 85.5%, MMLU-Pro: 80.5%.

What are the context window sizes for Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick?

Llama 3.1 Nemotron Ultra 253B v1 supports an unknown number of tokens and Llama 4 Maverick supports 1.0M tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick?

Key differences include LLM Stats Score (19.1 vs 14.7), multimodal support (no vs yes), licensing (Llama 3.1 Community License vs Llama 4 Community License Agreement). See the full comparison above for benchmark-by-benchmark results.

Who makes Llama 3.1 Nemotron Ultra 253B v1 and Llama 4 Maverick?

Llama 3.1 Nemotron Ultra 253B v1 is developed by NVIDIA and Llama 4 Maverick is developed by Meta.