- Organizations
- inclusionai
- Ling 3.0 Flash
Ling 3.0 Flash: API Pricing, Context Window & Benchmarks
Ling 3.0 Flash is a language model from Unknown Organization, released in August 2026, with a 131K-token context window, and pricing from $0.060/M input, $0.012/M cached input, $0.180/M output.
The model prioritizes token efficiency and agentic inference at production scale, stretching what developers can achieve within limited token, latency, and serving-cost budgets.
Ling 3.0 Flash benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Ling 3.0 Flash across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Ling 3.0 Flash holds up as conversations get longer.
Quality Tracker
Ling 3.0 Flash Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Ling 3.0 Flash pricing
Providers
Ling 3.0 Flash starts at $0.0600 per million input tokens and $0.180 per million output tokens via DeepInfra. Reused prompt prefixes cost $0.0120 per million cached input tokens.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0600 | $0.0120 | $0.180 | 131.1K/131.1K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Ling 3.0 Flash context window
Input and output token limits for Ling 3.0 Flash, plus how it ranks on long-context understanding.
Ling 3.0 Flash API
Available from the model provider
Ling 3.0 Flash is available from DeepInfra. It is not currently routed through the LLM Stats gateway.
Read the official API documentationLing 3.0 Flash latency
Ling 3.0 Flash time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Ling 3.0 Flash examples
Recent arena outputs from Ling 3.0 Flash, picked from the highest-ranked matchups.
Ling 3.0 Flash license
Ling 3.0 Flash has 124.0B parameters.
- Parameters
- 124.0B
Ling 3.0 Flash resources
Official sources for Ling 3.0 Flash: api documentation, model weights.
Ling 3.0 Flash vs other models
The most-compared alternatives to Ling 3.0 Flash are GPT OSS 120B High, Step-3.5-Flash, MiMo-V2-Pro. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Ling 3.0 Flash
Models ranked just above and below Ling 3.0 Flash by LLM Stats score.
FAQ
Common questions about Ling 3.0 Flash.