QwenReleased on Sep 19, 2024

Qwen2.5 72B Instruct: API Pricing, Context Window & Benchmarks

Qwen2.5 72B Instruct is a language model from Qwen, released in September 2024.

Qwen2.5-72B-Instruct is an instruction-tuned 72 billion parameter language model, part of the Qwen2.5 series. It is designed to follow instructions, generate long texts (over 8K tokens), understand structured data (e.g., tables), and

Input
Text
Output
Text

Qwen2.5 72B Instruct benchmarks

Rankings

Quality Tracker

Qwen2.5 72B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

Qwen2.5 72B Instruct pricing

Providers

Qwen2.5 72B Instruct starts at $0.350 per million input tokens and $0.400 per million output tokens via DeepInfra. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.350$0.400131.1K/8.2K
/0.50
10/
/
Hyperbolic logoHyperbolic
$0.400$0.400131.1K/8.2K
/0.50
100/
/
Fireworks logoFireworks
$0.890$0.890131.1K/8.2K
/0.37
59/
/
Together logoTogether
$1.20$1.20131.1K/8.2K
/0.50
47/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

Qwen2.5 72B Instruct context window

Input and output token limits for Qwen2.5 72B Instruct, plus how it ranks on long-context understanding.

InputOutput
131Ktokens
8Ktokens
197 pages of text
131K
8K128K1M

Qwen2.5 72B Instruct API

Available from the model provider

Qwen2.5 72B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Qwen2.5 72B Instruct latency

Qwen2.5 72B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Qwen2.5 72B Instruct examples

Recent arena outputs from Qwen2.5 72B Instruct, picked from the highest-ranked matchups.

Qwen2.5 72B Instruct license

Qwen2.5 72B Instruct is a proprietary model available under its provider's product and API terms, has 72.7B parameters.

License
Qwen
Hosted access
Parameters
72.7B

Alibaba Qwen License

Qwen2.5 72B Instruct resources

Official sources for Qwen2.5 72B Instruct: api documentation, official launch post, source repository, model weights.

Qwen2.5 72B Instruct vs other models

The most-compared alternatives to Qwen2.5 72B Instruct are Phi 4 Reasoning, DeepSeek R1 Distill Qwen 32B, Qwen2.5 32B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Qwen2.5 72B Instruct

Models ranked just above and below Qwen2.5 72B Instruct by LLM Stats score.

 

Phi 4 Reasoning

Score pending
 

DeepSeek R1 Distill Qwen 32B

Score pending
 

Qwen2.5 32B Instruct

Score pending
 

Phi 4

Score pending
 

DeepSeek R1 Distill Qwen 7B

Score pending
 

Qwen3 VL 8B Instruct

Score pending

FAQ

Common questions about Qwen2.5 72B Instruct.

When was Qwen2.5 72B Instruct released?

Qwen2.5 72B Instruct was released on September 19, 2024 by Qwen. This is the official Qwen2.5 72B Instruct release date tracked on LLM Stats.

How much does Qwen2.5 72B Instruct cost?

Qwen2.5 72B Instruct pricing starts at $0.35 per million input tokens and $0.40 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is Qwen2.5 72B Instruct available via API?

Yes, Qwen2.5 72B Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 4 providers tracked on LLM Stats.

How big is Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct has 72.7 billion parameters. It was trained on 18.0 trillion tokens.

Who created Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct was created by Qwen.

What is the license for Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is released under the Qwen license.

What is Qwen2.5 72B Instruct latency?

Qwen2.5 72B Instruct p95 time to first token is 0.37 seconds via Fireworks over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is available through 4 providers including DeepInfra, Hyperbolic, Fireworks, and 1 more.

Where is the Qwen2.5 72B Instruct paper or technical report?

Qwen2.5 72B Instruct has a paper or technical report available at https://qwenlm.github.io/blog/qwen2.5/. Use that source for architecture, training, release and evaluation details.

What models should I compare Qwen2.5 72B Instruct against?

Common Qwen2.5 72B Instruct comparisons include Qwen2.5 72B Instruct vs Phi 4 Reasoning, Qwen2.5 72B Instruct vs DeepSeek R1 Distill Qwen 32B, Qwen2.5 72B Instruct vs Qwen2.5 32B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.