OpenAIReleased on Aug 5, 2025

GPT OSS 120B High: API Pricing, Context Window & Benchmarks

GPT OSS 120B High is a language model from OpenAI, released in August 2025.

GPT-OSS-120B High provides enhanced reasoning capabilities with high-effort thinking for complex problems. This variant offers deeper analysis and more thorough responses compared to the base model, making it ideal for challenging tasks

Input
Text
Output
Text

GPT OSS 120B High benchmarks

Rankings

Quality Tracker

GPT OSS 120B High Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

GPT OSS 120B High pricing

Providers

GPT OSS 120B High starts at $0.100 per million input tokens and $0.500 per million output tokens via OpenAI. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
OpenAI logoOpenAI
$0.100$0.500131.1K/131.1K
/6.50
100/
/
Fireworks logoFireworks
$0.150$0.600131.0K/30.0K
0.00/0.00
/
0.00%(78)
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

GPT OSS 120B High model size

GPT OSS 120B High has 116.8 billion parameters. See how it compares to other models in the same parameter range.

Parameters
116.8B
Very large (80–200B)
116.8B
1B7B70B405B

GPT OSS 120B High context window

Input and output token limits for GPT OSS 120B High, plus how it ranks on long-context understanding.

InputOutput
131Ktokens
131Ktokens
197 pages of text
131K
8K128K1M

GPT OSS 120B High latency

GPT OSS 120B High time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Provider operational metrics

Time to first token, output throughput, and failed-request rate from live API traffic

Loading chart...
Loading chart...
Loading chart...

GPT OSS 120B High examples

Recent arena outputs from GPT OSS 120B High, picked from the highest-ranked matchups.

GPT OSS 120B High license

GPT OSS 120B High is released under the Apache 2.0 license, which permits commercial use, has 116.8B parameters.

License
Apache 2.0
Commercial use allowed
Parameters
116.8B

Apache License 2.0 - allows commercial use

GPT OSS 120B High resources

Official sources for GPT OSS 120B High: official playground, paper or system card, official launch post, source repository, model weights.

GPT OSS 120B High vs other models

The most-compared alternatives to GPT OSS 120B High are Seed 2.0 Lite, K-EXAONE-236B-A23B, LongCat-Flash-Thinking-2601. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like GPT OSS 120B High

Models ranked just above and below GPT OSS 120B High by LLM Stats score.

 

Seed 2.0 Lite

Score pending
 

K-EXAONE-236B-A23B

Score pending
 

LongCat-Flash-Thinking-2601

Score pending
 

Nova 2 Pro

Score pending
 

Claude Opus 4.1

Score pending
 

Nova 2 Omni

Score pending

FAQ

Common questions about GPT OSS 120B High.

When was GPT OSS 120B High released?

GPT OSS 120B High was released on August 5, 2025 by OpenAI. This is the official GPT OSS 120B High release date tracked on LLM Stats.

How much does GPT OSS 120B High cost?

GPT OSS 120B High pricing starts at $0.10 per million input tokens and $0.50 per million output tokens via OpenAI, the lowest price among tracked providers.

How big is GPT OSS 120B High?

GPT OSS 120B High has 116.8 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created GPT OSS 120B High?

GPT OSS 120B High was created by OpenAI.

What is the license for GPT OSS 120B High?

GPT OSS 120B High is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

What is GPT OSS 120B High latency?

GPT OSS 120B High p95 time to first token is 12.14 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use GPT OSS 120B High?

GPT OSS 120B High is available through 2 providers including OpenAI, Fireworks.

Where is the GPT OSS 120B High paper or technical report?

GPT OSS 120B High has a paper or technical report available at https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare GPT OSS 120B High against?

Common GPT OSS 120B High comparisons include GPT OSS 120B High vs Seed 2.0 Lite, GPT OSS 120B High vs K-EXAONE-236B-A23B, GPT OSS 120B High vs LongCat-Flash-Thinking-2601. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.