The AI arena is free today

Open Superagent

DeepSeek-V3.2 (Thinking) vs Llama 3.1 Nemotron Ultra 253B v1

DeepSeek-V3.2 (Thinking) leads the LLM Stats Score 32.9 to 19.3.

DeepSeek · NVIDIA · Updated for 2026

Which is better?

DeepSeek-V3.2 (Thinking) leads the overall LLM Stats Score 32.9 to 19.3, ranking #96 overall.

In the 3 individual benchmarks reported for both models, DeepSeek-V3.2 (Thinking) wins 3; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose DeepSeek-V3.2 (Thinking)

  • overall performance matters — it scores 32.9 and ranks #96 on LLM Stats
  • your work emphasizes reasoning — it leads those capability indexes
  • you value its reported benchmark strengths — it wins 3 of 3 exact shared results
  • you want the most recent training data — it shipped Dec 2025

Choose Llama 3.1 Nemotron Ultra 253B v1

  • you are already invested in the NVIDIA ecosystem

At a glance

The differences that matter most.

Core performance indexes
32.9
#96
19.3
#196
32.9
#95
18.9
#194
22.9
#75
12.6
#140
Cost, coverage & limits
Benchmark wins
3 of 3
0 of 3
Input price
$0.28 / M
— / M
Output price
$0.42 / M
— / M
Context window
131,072

Capability indexes

Additional strengths measured across groups of related public benchmarks

1 shared
Index
DeepSeek-V3.2 (Thinking)
Llama 3.1 Nemotron Ultra 253B v1
30.6#71
15.7#207
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

14 reported for DeepSeek-V3.2 (Thinking) · 6 for Llama 3.1 Nemotron Ultra 253B v1

3 shared

DeepSeek-V3.2 (Thinking) outperforms in 3 benchmarks (AIME 2025, GPQA, LiveCodeBench), while Llama 3.1 Nemotron Ultra 253B v1 is better at 0 benchmarks.

DeepSeek-V3.2 (Thinking) significantly outperforms across most benchmarks.

Tue Sep 08 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Model Size

Parameter count comparison

432.0B diff

DeepSeek-V3.2 (Thinking) has 432.0B more parameters than Llama 3.1 Nemotron Ultra 253B v1, making it 170.8% larger.

DeepSeek
DeepSeek-V3.2 (Thinking)
685.0Bparameters
NVIDIA
Llama 3.1 Nemotron Ultra 253B v1
253.0Bparameters
685.0B
DeepSeek-V3.2 (Thinking)
253.0B
Llama 3.1 Nemotron Ultra 253B v1

Context Window

Maximum input and output token capacity

Only DeepSeek-V3.2 (Thinking) specifies input context (131,072 tokens). Only DeepSeek-V3.2 (Thinking) specifies output context (65,536 tokens).

DeepSeek
DeepSeek-V3.2 (Thinking)
Input131,072 tokens
Output65,536 tokens
NVIDIA
Llama 3.1 Nemotron Ultra 253B v1
Input- tokens
Output- tokens
Tue Sep 08 2026 • llm-stats.com

License

Usage and distribution terms

DeepSeek-V3.2 (Thinking) is licensed under MIT, while Llama 3.1 Nemotron Ultra 253B v1 uses Llama 3.1 Community License.

License differences may affect how you can use these models in commercial or open-source projects.

DeepSeek-V3.2 (Thinking)

MIT

Open weights

Llama 3.1 Nemotron Ultra 253B v1

Llama 3.1 Community License

Open weights

Release Timeline

When each model was launched

DeepSeek-V3.2 (Thinking) was released on 2025-12-01, while Llama 3.1 Nemotron Ultra 253B v1 was released on 2025-04-07.

DeepSeek-V3.2 (Thinking) is 8 months newer than Llama 3.1 Nemotron Ultra 253B v1.

DeepSeek-V3.2 (Thinking)

Dec 1, 2025

9 months ago

7mo newer
Llama 3.1 Nemotron Ultra 253B v1

Apr 7, 2025

1.4 years ago

Knowledge Cutoff

When training data ends

Llama 3.1 Nemotron Ultra 253B v1 has a documented knowledge cutoff of 2023-12-01, while DeepSeek-V3.2 (Thinking)'s cutoff date is not specified.

We can confirm Llama 3.1 Nemotron Ultra 253B v1's training data extends to 2023-12-01, but cannot make a direct comparison without DeepSeek-V3.2 (Thinking)'s cutoff date.

DeepSeek-V3.2 (Thinking)

Llama 3.1 Nemotron Ultra 253B v1

Dec 2023

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against DeepSeek-V3.2 (Thinking) and Llama 3.1 Nemotron Ultra 253B v1 side-by-side, then vote on the output you prefer.

DeepSeek-V3.2 (Thinking)
✓ Preferred
Llama 3.1 Nemotron Ultra 253B v1
Open in Playground

FAQ

Common questions about DeepSeek-V3.2 (Thinking) vs Llama 3.1 Nemotron Ultra 253B v1.

Which is better, DeepSeek-V3.2 (Thinking) or Llama 3.1 Nemotron Ultra 253B v1?

DeepSeek-V3.2 (Thinking) leads the LLM Stats Score 32.9 to 19.3. DeepSeek-V3.2 (Thinking) is made by DeepSeek and Llama 3.1 Nemotron Ultra 253B v1 is made by NVIDIA. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does DeepSeek-V3.2 (Thinking) compare to Llama 3.1 Nemotron Ultra 253B v1 in benchmarks?

DeepSeek-V3.2 (Thinking) scores AIME 2025: 93.1%, HMMT 2025: 90.2%, MMLU-Pro: 85.0%, LiveCodeBench: 83.3%, GPQA: 82.4%. Llama 3.1 Nemotron Ultra 253B v1 scores MATH-500: 97.0%, IFEval: 89.5%, GPQA: 76.0%, BFCL v2: 74.1%, AIME 2025: 72.5%.

What are the context window sizes for DeepSeek-V3.2 (Thinking) and Llama 3.1 Nemotron Ultra 253B v1?

DeepSeek-V3.2 (Thinking) supports 131K tokens and Llama 3.1 Nemotron Ultra 253B v1 supports an unknown number of tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between DeepSeek-V3.2 (Thinking) and Llama 3.1 Nemotron Ultra 253B v1?

Key differences include LLM Stats Score (32.9 vs 19.3), licensing (MIT vs Llama 3.1 Community License). See the full comparison above for benchmark-by-benchmark results.

Who makes DeepSeek-V3.2 (Thinking) and Llama 3.1 Nemotron Ultra 253B v1?

DeepSeek-V3.2 (Thinking) is developed by DeepSeek and Llama 3.1 Nemotron Ultra 253B v1 is developed by NVIDIA.