- Organizations
- Qwen
- Qwen3 Max
Qwen3 Max: Benchmarks, Pricing & Context Window
Qwen3 Max is a language model from Qwen, released in December 2025, with a 256K-token context window, and pricing from $1.20/M input, $0.240/M cached input, $6.00/M output.
Qwen3 Max is the flagship model in the Qwen3 series, designed for maximum performance across all tasks. It excels in reasoning, mathematics, coding, and complex problem-solving while maintaining strong multilingual capabilities. Optimized
Qwen3 Max benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Qwen3 Max across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Qwen3 Max holds up as conversations get longer.
Quality Tracker
Qwen3 Max Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Qwen3 Max pricing
Providers
Qwen3 Max starts at $0.500 per million input tokens and $5.00 per million output tokens via Novita. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.500 | — | $5.00 | 256.0K/131.1K | — | — | / | |
| $1.20 | $0.240 | $6.00 | 256.0K/256.0K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Qwen3 Max context window
Input and output token limits for Qwen3 Max, plus how it ranks on long-context understanding.
Try now
Make it with
Qwen3 Max.
Qwen3 Max latency
Qwen3 Max time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Qwen3 Max examples
Recent arena outputs from Qwen3 Max, picked from the highest-ranked matchups.
Qwen3 Max license
Qwen3 Max is a proprietary model available under its provider's product and API terms, has 1.0T parameters.
- License
- Proprietary
- Hosted access
- Parameters
- 1.0T
Proprietary license - usage restrictions apply
Qwen3 Max resources
Official sources for Qwen3 Max: provider documentation, official playground, paper or system card, source repository.
Qwen3 Max vs other models
The most-compared alternatives to Qwen3 Max are Claude 3.5 Sonnet, LongCat-Flash-Thinking-2601, Nova 2 Pro. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Qwen3 Max
Models ranked just above and below Qwen3 Max by LLM Stats score.
FAQ
Common questions about Qwen3 Max.