- Organizations
- OpenAI
- GPT OSS 120B
GPT OSS 120B: Benchmarks, Pricing & Context Window
GPT OSS 120B is a language model from OpenAI, released in August 2025, with a 131K-token context window, and pricing from $0.037/M input and $0.170/M output.
GPT-OSS-120B is an open-weight, 116.8B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is
GPT OSS 120B benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for GPT OSS 120B across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How GPT OSS 120B holds up as conversations get longer.
Quality Tracker
GPT OSS 120B Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
GPT OSS 120B pricing
Providers
GPT OSS 120B starts at $0.0370 per million input tokens and $0.170 per million output tokens via DeepInfra. See all 5 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0370 | — | $0.170 | 131.1K/131.1K | 0.00 | 1 | / | |
| $0.100 | — | $0.500 | 131.1K/131.1K | — | — | / | |
| $0.100 | — | $0.500 | 131.1K/131.1K | 5.20 | — | / | |
| $0.150 | — | $0.600 | 131.0K/30.0K | 0.00 | 4 | / | |
| $0.150 | — | $0.600 | 131.0K/30.0K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
GPT OSS 120B model size
GPT OSS 120B has 116.8 billion parameters. See how it compares to other models in the same parameter range.
GPT OSS 120B context window
Input and output token limits for GPT OSS 120B, plus how it ranks on long-context understanding.
Try now
Make it with
GPT OSS 120B.
GPT OSS 120B latency
GPT OSS 120B time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Provider operational metrics
Time to first token, output throughput, and failed-request rate from live model usage
GPT OSS 120B examples
Recent arena outputs from GPT OSS 120B, picked from the highest-ranked matchups.
GPT OSS 120B license
GPT OSS 120B is released under the Apache 2.0 license, which permits commercial use, has 116.8B parameters.
- License
- Apache 2.0
- Commercial use allowed
- Parameters
- 116.8B
Apache License 2.0 - allows commercial use
GPT OSS 120B resources
Official sources for GPT OSS 120B: official playground, paper or system card, official launch post, source repository, model weights.
GPT OSS 120B vs other models
The most-compared alternatives to GPT OSS 120B are Kimi K2 0905, EXAONE 4.5 33B, LongCat-Flash-Thinking-2601. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like GPT OSS 120B
Models ranked just above and below GPT OSS 120B by LLM Stats score.
FAQ
Common questions about GPT OSS 120B.