- Organizations
- Qwen
- Qwen3.8 Flash
Qwen3.8 Flash: API Pricing, Context Window & Benchmarks
Qwen3.8 Flash is a language model from Qwen, released in August 2026, with multimodal input, a 1M-token context window, and pricing from $0.150/M input, $0.016/M cached input, $0.470/M output.
Qwen3.8 Flash is the production QwenCloud / OpenRouter API model (id qwen3.8-flash), not the open-weight Qwen3.8-Flash-Next checkpoint. Hugging Face states Flash is the official managed version based on Flash-Next with production features
Qwen3.8 Flash benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Qwen3.8 Flash across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Qwen3.8 Flash holds up as conversations get longer.
Quality Tracker
Qwen3.8 Flash Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Qwen3.8 Flash pricing
Providers
Qwen3.8 Flash starts at $0.150 per million input tokens and $0.470 per million output tokens via Novita. Reused prompt prefixes cost $0.0160 per million cached input tokens.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.150 | $0.0160 | $0.470 | 1.0M/131.1K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Qwen3.8 Flash context window
Input and output token limits for Qwen3.8 Flash, plus how it ranks on long-context understanding.
Qwen3.8 Flash API
Run a request to see the response
Use it in your code
Billed at $0.15 input / $0.02 cached input / $0.47 output per 1M tokens through the LLM Stats gateway.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://gateway.llm-stats.com/v1"
)
response = client.chat.completions.create(
model="qwen3.8-flash",
messages=[
{"role": "user", "content": "What is machine learning?"}
]
)
print(response.choices[0].message.content)Need an API key? Create one above in the playground, or read the API documentation.
Qwen3.8 Flash latency
Qwen3.8 Flash time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Qwen3.8 Flash examples
Recent arena outputs from Qwen3.8 Flash, picked from the highest-ranked matchups.
Qwen3.8 Flash license
Qwen3.8 Flash is a proprietary model available under its provider's product and API terms, has 125.0B parameters.
- License
- Proprietary
- Hosted access
- Parameters
- 125.0B
Proprietary license - usage restrictions apply
Qwen3.8 Flash resources
Official sources for Qwen3.8 Flash: api documentation, official playground, official launch post, model weights.
Qwen3.8 Flash vs other models
The most-compared alternatives to Qwen3.8 Flash are Claude Opus 4.6, Gemini 3 Pro, GPT-5.2. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Qwen3.8 Flash
Models ranked just above and below Qwen3.8 Flash by LLM Stats score.
FAQ
Common questions about Qwen3.8 Flash.