OpenAIReleased on Jan 30, 2025

o3-mini: API Pricing, Context Window & Benchmarks

o3-mini is a language model from OpenAI, released in January 2025.

A smaller variant of O3, expected to offer enhanced multimodal capabilities, improved reasoning, and more efficient resource utilization compared to previous models while maintaining strong performance on core tasks.

Input
Text
Output
Text

o3-mini benchmarks

Rankings

Quality Tracker

o3-mini Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Jul 29 2026
Notice missing or incorrect data?

o3-mini pricing

Providers

o3-mini starts at $1.10 per million input tokens and $4.40 per million output tokens via Azure. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Azure logoAzure
$1.10$4.40200.0K/100.0K
/5.20
115/
/
OpenAI logoOpenAI
$1.10$4.40200.0K/100.0K
/5.20
115/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

o3-mini context window

Input and output token limits for o3-mini, plus how it ranks on long-context understanding.

InputOutput
200Ktokens
100Ktokens
301 pages of text
200K
8K128K1M

o3-mini API

Available from the model provider

o3-mini has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

o3-mini latency

o3-mini time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

o3-mini examples

Recent arena outputs from o3-mini, picked from the highest-ranked matchups.

o3-mini license

o3-mini is a proprietary model available under its provider's product and API terms, has a knowledge cutoff of September 2023.

License
Proprietary
Hosted access
Knowledge cutoff
September 2023

Proprietary license - usage restrictions apply

o3-mini resources

Official sources for o3-mini: api documentation, paper or system card, official launch post, source repository.

o3-mini vs other models

The most-compared alternatives to o3-mini are Kimi-k1.5, Claude 3 Opus, Llama 3.1 Nemotron Ultra 253B v1. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like o3-mini

Models ranked just above and below o3-mini by LLM Stats score.

 

Kimi-k1.5

Score pending
 

Claude 3 Opus

Score pending
 

Llama 3.1 Nemotron Ultra 253B v1

Score pending
 

Llama 3.1 405B Instruct

Score pending
 

GPT-4 Turbo

Score pending
 

Claude 3.5 Sonnet

Score pending

FAQ

Common questions about o3-mini.

When was o3-mini released?

o3-mini was released on January 30, 2025 by OpenAI. This is the official o3-mini release date tracked on LLM Stats.

How much does o3-mini cost?

o3-mini pricing starts at $1.10 per million input tokens and $4.40 per million output tokens via Azure, the lowest price among tracked providers.

Is o3-mini available via API?

Yes, o3-mini is available via API. See the official documentation for authentication and endpoint details. It is served by 2 providers tracked on LLM Stats.

Who created o3-mini?

o3-mini was created by OpenAI.

What is the license for o3-mini?

o3-mini is released under the Proprietary license.

What is the knowledge cutoff date for o3-mini?

o3-mini has a knowledge cutoff of September 2023, meaning it was trained on data up to that point and may not know about events after it.

What is o3-mini latency?

o3-mini p95 time to first token is 5.20 seconds via Azure over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use o3-mini?

o3-mini is available through 2 providers including Azure, OpenAI.

Where is the o3-mini paper or technical report?

o3-mini has a paper or technical report available at https://cdn.openai.com/o3-mini-system-card.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare o3-mini against?

Common o3-mini comparisons include o3-mini vs Kimi-k1.5, o3-mini vs Claude 3 Opus, o3-mini vs Llama 3.1 Nemotron Ultra 253B v1. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.