- Organizations
- Qwen
- Qwen2.5 72B Instruct
Qwen2.5 72B Instruct: API Pricing, Context Window & Benchmarks
Qwen2.5 72B Instruct is a language model from Qwen, released in September 2024.
Qwen2.5-72B-Instruct is an instruction-tuned 72 billion parameter language model, part of the Qwen2.5 series. It is designed to follow instructions, generate long texts (over 8K tokens), understand structured data (e.g., tables), and
Qwen2.5 72B Instruct benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Qwen2.5 72B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Qwen2.5 72B Instruct holds up as conversations get longer.
Quality Tracker
Qwen2.5 72B Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Qwen2.5 72B Instruct pricing
Providers
Qwen2.5 72B Instruct starts at $0.350 per million input tokens and $0.400 per million output tokens via DeepInfra. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.350 | — | $0.400 | 131.1K/8.2K | 0.50 | — | / | |
| $0.400 | — | $0.400 | 131.1K/8.2K | 0.50 | — | / | |
| $0.890 | — | $0.890 | 131.1K/8.2K | 0.37 | — | / | |
| $1.20 | — | $1.20 | 131.1K/8.2K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Qwen2.5 72B Instruct context window
Input and output token limits for Qwen2.5 72B Instruct, plus how it ranks on long-context understanding.
Qwen2.5 72B Instruct API
Available from the model provider
Qwen2.5 72B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.
Read the official API documentationQwen2.5 72B Instruct latency
Qwen2.5 72B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Qwen2.5 72B Instruct examples
Recent arena outputs from Qwen2.5 72B Instruct, picked from the highest-ranked matchups.
Qwen2.5 72B Instruct license
Qwen2.5 72B Instruct is a proprietary model available under its provider's product and API terms, has 72.7B parameters.
- License
- Qwen
- Hosted access
- Parameters
- 72.7B
Alibaba Qwen License
Qwen2.5 72B Instruct resources
Official sources for Qwen2.5 72B Instruct: api documentation, official launch post, source repository, model weights.
Qwen2.5 72B Instruct vs other models
The most-compared alternatives to Qwen2.5 72B Instruct are GPT-4 Turbo, Phi 4 Reasoning, DeepSeek R1 Distill Qwen 32B. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Qwen2.5 72B Instruct
Models ranked just above and below Qwen2.5 72B Instruct by LLM Stats score.
FAQ
Common questions about Qwen2.5 72B Instruct.