- Organizations
- OpenAI
- GPT OSS 20B
GPT OSS 20B: API Pricing, Context Window & Benchmarks
GPT OSS 20B is a language model from OpenAI, released in August 2025.
The gpt-oss-20b model (technically 20.9B parameters) achieves near-parity with OpenAI o4-mini on core reasoning benchmarks, while running efficiently on a single 80 GB GPU. The gpt-oss-20b model delivers similar results to OpenAI o3‑mini
GPT OSS 20B benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for GPT OSS 20B across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How GPT OSS 20B holds up as conversations get longer.
Quality Tracker
GPT OSS 20B Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
GPT OSS 20B pricing
Providers
GPT OSS 20B starts at $0.0500 per million input tokens and $0.200 per million output tokens via Novita. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0500 | — | $0.200 | 131.1K/32.8K | — | — | / | |
| $0.100 | — | $0.500 | 131.0K/30.0K | — | — | / | |
| $0.100 | — | $0.500 | 131.0K/30.0K | 0.38 | — | / | |
| $0.100 | — | $0.500 | 131.1K/131.1K | 5.20 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
GPT OSS 20B model size
GPT OSS 20B has 20.9 billion parameters. See how it compares to other models in the same parameter range.
GPT OSS 20B context window
Input and output token limits for GPT OSS 20B, plus how it ranks on long-context understanding.
GPT OSS 20B latency
GPT OSS 20B time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
GPT OSS 20B examples
Recent arena outputs from GPT OSS 20B, picked from the highest-ranked matchups.
GPT OSS 20B license
GPT OSS 20B is released under the Apache 2.0 license, which permits commercial use, has 20.9B parameters.
- License
- Apache 2.0
- Commercial use allowed
- Parameters
- 20.9B
Apache License 2.0 - allows commercial use
GPT OSS 20B resources
Official sources for GPT OSS 20B: official playground, paper or system card, official launch post, source repository, model weights.
GPT OSS 20B vs other models
The most-compared alternatives to GPT OSS 20B are Sarvam-105B, o1-mini, Qwen3 VL 8B Thinking. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like GPT OSS 20B
Models ranked just above and below GPT OSS 20B by LLM Stats score.
FAQ
Common questions about GPT OSS 20B.