QwenReleased on Sep 19, 2024

Qwen2.5-Coder 32B Instruct: API Pricing, Context Window & Benchmarks

Qwen2.5-Coder 32B Instruct is a language model from Qwen, released in September 2024.

Qwen2.5-Coder is a specialized coding model trained on 5.5 trillion tokens of code data, supporting 92 programming languages with a 128K context window. It excels in code generation, completion, repair, and multi-programming tasks while

Input
Text
Output
Text

Qwen2.5-Coder 32B Instruct benchmarks

Rankings

Quality Tracker

Qwen2.5-Coder 32B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Jul 22 2026
Notice missing or incorrect data?

Qwen2.5-Coder 32B Instruct pricing

Providers

Qwen2.5-Coder 32B Instruct starts at $0.0900 per million input tokens and $0.0900 per million output tokens via Lambda. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Lambda logoLambda
$0.0900$0.0900128.0K/128.0K
/0.50
42/
/
DeepInfra logoDeepInfra
$0.180$0.180128.0K/128.0K
/0.50
44/
/
Hyperbolic logoHyperbolic
$0.200$0.200128.0K/128.0K
/0.50
100/
/
Fireworks logoFireworks
$0.890$0.890128.0K/128.0K
/0.26
110/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

Qwen2.5-Coder 32B Instruct model size

Qwen2.5-Coder 32B Instruct has 32 billion parameters and was trained on 5.5 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
32B
5.5Ttokens
172× tokens-to-params ratio
Large (30–80B)
32B
1B7B70B405B

Qwen2.5-Coder 32B Instruct context window

Input and output token limits for Qwen2.5-Coder 32B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Qwen2.5-Coder 32B Instruct API

Available from the model provider

Qwen2.5-Coder 32B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Qwen2.5-Coder 32B Instruct latency

Qwen2.5-Coder 32B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Qwen2.5-Coder 32B Instruct examples

Recent arena outputs from Qwen2.5-Coder 32B Instruct, picked from the highest-ranked matchups.

Qwen2.5-Coder 32B Instruct license

Qwen2.5-Coder 32B Instruct is released under the Apache 2.0 license, which permits commercial use, has 32.0B parameters.

License
Apache 2.0
Commercial use allowed
Parameters
32.0B

Apache License 2.0 - allows commercial use

Qwen2.5-Coder 32B Instruct resources

Official sources for Qwen2.5-Coder 32B Instruct: api documentation, paper or system card, official launch post, source repository, model weights.

Qwen2.5-Coder 32B Instruct vs other models

The most-compared alternatives to Qwen2.5-Coder 32B Instruct are Claude 3 Haiku, Gemma 2 27B, Llama 3.2 11B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Qwen2.5-Coder 32B Instruct

Models ranked just above and below Qwen2.5-Coder 32B Instruct by LLM Stats score.

 

Claude 3 Haiku

Score pending
 

Gemma 2 27B

Score pending
 

Llama 3.2 11B Instruct

Score pending
 

Phi-3.5-MoE-instruct

Score pending
 

Llama 4 Scout

Score pending
 

Llama 3.1 8B Instruct

Score pending

FAQ

Common questions about Qwen2.5-Coder 32B Instruct.

When was Qwen2.5-Coder 32B Instruct released?

Qwen2.5-Coder 32B Instruct was released on September 19, 2024 by Qwen. This is the official Qwen2.5-Coder 32B Instruct release date tracked on LLM Stats.

How much does Qwen2.5-Coder 32B Instruct cost?

Qwen2.5-Coder 32B Instruct pricing starts at $0.09 per million input tokens and $0.09 per million output tokens via Lambda, the lowest price among tracked providers.

Is Qwen2.5-Coder 32B Instruct available via API?

Yes, Qwen2.5-Coder 32B Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 4 providers tracked on LLM Stats.

How big is Qwen2.5-Coder 32B Instruct?

Qwen2.5-Coder 32B Instruct has 32 billion parameters. It was trained on 5.5 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Qwen2.5-Coder 32B Instruct?

Qwen2.5-Coder 32B Instruct was created by Qwen.

What is the license for Qwen2.5-Coder 32B Instruct?

Qwen2.5-Coder 32B Instruct is released under the Apache 2.0 license. This is an open-source / open-weight license that permits self-hosting.

What is Qwen2.5-Coder 32B Instruct latency?

Qwen2.5-Coder 32B Instruct p95 time to first token is 0.26 seconds via Fireworks over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Qwen2.5-Coder 32B Instruct?

Qwen2.5-Coder 32B Instruct is available through 4 providers including Lambda, DeepInfra, Hyperbolic, and 1 more.

Where is the Qwen2.5-Coder 32B Instruct paper or technical report?

Qwen2.5-Coder 32B Instruct has a paper or technical report available at https://arxiv.org/abs/2409.12186. Use that source for architecture, training, release and evaluation details.

What models should I compare Qwen2.5-Coder 32B Instruct against?

Common Qwen2.5-Coder 32B Instruct comparisons include Qwen2.5-Coder 32B Instruct vs Claude 3 Haiku, Qwen2.5-Coder 32B Instruct vs Gemma 2 27B, Qwen2.5-Coder 32B Instruct vs Llama 3.2 11B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.