The AI arena is free today

Open Superagent
OpenAIReleased on Aug 5, 2025

GPT OSS 120B: Benchmarks, Pricing & Context Window

GPT OSS 120B is a language model from OpenAI, released in August 2025, with a 131K-token context window, and pricing from $0.037/M input and $0.170/M output.

GPT-OSS-120B is an open-weight, 116.8B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is

Input
Text
Output
Text

GPT OSS 120B benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for GPT OSS 120B across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How GPT OSS 120B holds up as conversations get longer.

Quality Tracker

GPT OSS 120B Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sun Sep 13 2026
Notice missing or incorrect data?

GPT OSS 120B pricing

Providers

GPT OSS 120B starts at $0.0370 per million input tokens and $0.170 per million output tokens via DeepInfra. See all 5 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0370$0.170131.1K/131.1K
0.00
1
/
Novita logoNovita
$0.100$0.500131.1K/131.1K
/
OpenAI logoOpenAI
$0.100$0.500131.1K/131.1K
5.20
/
Fireworks logoFireworks
$0.150$0.600131.0K/30.0K
0.00
4
/
Groq logoGroq
$0.150$0.600131.0K/30.0K
0.50
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

GPT OSS 120B model size

GPT OSS 120B has 116.8 billion parameters. See how it compares to other models in the same parameter range.

Parameters
116.8B
Very large (80–200B)
116.8B
1B7B70B405B

GPT OSS 120B context window

Input and output token limits for GPT OSS 120B, plus how it ranks on long-context understanding.

InputOutput
131Ktokens
131Ktokens
197 pages of text
131K
8K128K1M

Try now

huggle
GPT OSS 120Bin Huggle

Make it with
GPT OSS 120B.

GPT OSS 120B

GPT OSS 120B latency

GPT OSS 120B time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Provider operational metrics

Time to first token, output throughput, and failed-request rate from live model usage

Loading chart...
Loading chart...
Loading chart...

GPT OSS 120B examples

Recent arena outputs from GPT OSS 120B, picked from the highest-ranked matchups.

GPT OSS 120B license

GPT OSS 120B is released under the Apache 2.0 license, which permits commercial use, has 116.8B parameters.

License
Apache 2.0
Commercial use allowed
Parameters
116.8B

Apache License 2.0 - allows commercial use

GPT OSS 120B resources

Official sources for GPT OSS 120B: official playground, paper or system card, official launch post, source repository, model weights.

GPT OSS 120B vs other models

The most-compared alternatives to GPT OSS 120B are Kimi K2 0905, EXAONE 4.5 33B, LongCat-Flash-Thinking-2601. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like GPT OSS 120B

Models ranked just above and below GPT OSS 120B by LLM Stats score.

 

Kimi K2 0905

Score pending
 

EXAONE 4.5 33B

Score pending
 

LongCat-Flash-Thinking-2601

Score pending
 

LongCat-Flash-Chat

Score pending
 

Qwen3-235B-A22B-Thinking-2507

Score pending
 

Nova 2 Lite

Score pending

FAQ

Common questions about GPT OSS 120B.

When was GPT OSS 120B released?

GPT OSS 120B was released on August 5, 2025 by OpenAI. This is the official GPT OSS 120B release date tracked on LLM Stats.

How much does GPT OSS 120B cost?

GPT OSS 120B pricing starts at $0.04 per million input tokens and $0.17 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is GPT OSS 120B?

GPT OSS 120B has 116.8 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created GPT OSS 120B?

GPT OSS 120B was created by OpenAI.

What is the license for GPT OSS 120B?

GPT OSS 120B is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

What is GPT OSS 120B latency?

GPT OSS 120B p95 time to first token is 0.50 seconds via Groq over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use GPT OSS 120B?

GPT OSS 120B is available through 5 providers including DeepInfra, Novita, OpenAI, and 2 more.

Where is the GPT OSS 120B paper or technical report?

GPT OSS 120B has a paper or technical report available at https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare GPT OSS 120B against?

Common GPT OSS 120B comparisons include GPT OSS 120B vs Kimi K2 0905, GPT OSS 120B vs EXAONE 4.5 33B, GPT OSS 120B vs LongCat-Flash-Thinking-2601. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.