The AI arena is free today

Open Superagent

EXAONE 4.5 33B vs GPT OSS 120B

EXAONE 4.5 33B and GPT OSS 120B are closely matched at 26.0 and 28.8 on the LLM Stats Score.

LG AI Research · OpenAI · Updated for 2026

Which is better?

EXAONE 4.5 33B and GPT OSS 120B are closely matched on the overall LLM Stats Score at 26.0 and 28.8.

In the 1 individual benchmarks reported for both models, EXAONE 4.5 33B wins 1; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose EXAONE 4.5 33B

  • you value its reported benchmark strengths — it wins 1 of 1 exact shared results
  • you want the most recent training data — it shipped Apr 2026

Choose GPT OSS 120B

  • you need open weights you can self-host or fine-tune

At a glance

The differences that matter most.

Core performance indexes
26.0
#151
28.8
#131
26.6
#144
23.0
#165
Cost, coverage & limits
Benchmark wins
1 of 1
0 of 1
Input price
— / M
$0.04 / M
Output price
— / M
$0.17 / M
Context window
131,072

Capability indexes

Additional strengths measured across groups of related public benchmarks

1 shared
Index
EXAONE 4.5 33B
GPT OSS 120B
28.4#93
23.3#131
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

15 reported for EXAONE 4.5 33B · 7 for GPT OSS 120B

1 shared

EXAONE 4.5 33B outperforms in 1 benchmarks (GPQA), while GPT OSS 120B is better at 0 benchmarks.

EXAONE 4.5 33B significantly outperforms across most benchmarks.

Sun Sep 13 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Model Size

Parameter count comparison

83.8B diff

GPT OSS 120B has 83.8B more parameters than EXAONE 4.5 33B, making it 253.9% larger.

LG AI Research
EXAONE 4.5 33B
33.0Bparameters
OpenAI
GPT OSS 120B
116.8Bparameters
33.0B
EXAONE 4.5 33B
116.8B
GPT OSS 120B

Context Window

Maximum input and output token capacity

Only GPT OSS 120B specifies input context (131,072 tokens). Only GPT OSS 120B specifies output context (131,072 tokens).

LG AI Research
EXAONE 4.5 33B
Input- tokens
Output- tokens
OpenAI
GPT OSS 120B
Input131,072 tokens
Output131,072 tokens
Sun Sep 13 2026 • llm-stats.com

Input capabilities

Documented input modalities across available providers

EXAONE 4.5 33B supports multimodal inputs, whereas GPT OSS 120B does not.

EXAONE 4.5 33B can handle both text and other forms of data like images, making it suitable for multimodal applications.

EXAONE 4.5 33B

Text
Images
Audio
Video

GPT OSS 120B

Text
Images
Audio
Video

License

Usage and distribution terms

EXAONE 4.5 33B is licensed under a proprietary license, while GPT OSS 120B uses Apache 2.0.

License differences may affect how you can use these models in commercial or open-source projects.

EXAONE 4.5 33B

Proprietary

Closed source

GPT OSS 120B

Apache 2.0

Open weights

Release Timeline

When each model was launched

EXAONE 4.5 33B was released on 2026-04-09, while GPT OSS 120B was released on 2025-08-05.

EXAONE 4.5 33B is 8 months newer than GPT OSS 120B.

EXAONE 4.5 33B

Apr 9, 2026

5 months ago

8mo newer
GPT OSS 120B

Aug 5, 2025

1.1 years ago

Knowledge Cutoff

When training data ends

EXAONE 4.5 33B has a documented knowledge cutoff of 2024-12-01, while GPT OSS 120B's cutoff date is not specified.

We can confirm EXAONE 4.5 33B's training data extends to 2024-12-01, but cannot make a direct comparison without GPT OSS 120B's cutoff date.

EXAONE 4.5 33B

Dec 2024

GPT OSS 120B

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against EXAONE 4.5 33B and GPT OSS 120B side-by-side, then vote on the output you prefer.

EXAONE 4.5 33B
✓ Preferred
GPT OSS 120B
Open in Playground

FAQ

Common questions about EXAONE 4.5 33B vs GPT OSS 120B.

Which is better, EXAONE 4.5 33B or GPT OSS 120B?

EXAONE 4.5 33B and GPT OSS 120B are closely matched on the LLM Stats Score at 26.0 and 28.8. EXAONE 4.5 33B is made by LG AI Research and GPT OSS 120B is made by OpenAI. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does EXAONE 4.5 33B compare to GPT OSS 120B in benchmarks?

EXAONE 4.5 33B scores AIME 2025: 92.9%, AIME 2026: 92.6%, IFEval: 89.6%, AI2D: 89.0%, MathVista-Mini: 85.0%. GPT OSS 120B scores MMLU: 90.0%, CodeForces: 82.1%, GPQA: 80.1%, TAU-bench Retail: 67.8%, HealthBench: 57.6%.

What are the context window sizes for EXAONE 4.5 33B and GPT OSS 120B?

EXAONE 4.5 33B supports an unknown number of tokens and GPT OSS 120B supports 131K tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between EXAONE 4.5 33B and GPT OSS 120B?

Key differences include LLM Stats Score (26.0 vs 28.8), multimodal support (yes vs no), licensing (Proprietary vs Apache 2.0). See the full comparison above for benchmark-by-benchmark results.

Who makes EXAONE 4.5 33B and GPT OSS 120B?

EXAONE 4.5 33B is developed by LG AI Research and GPT OSS 120B is developed by OpenAI.