- Organizations
- MoonshotAI
- Kimi K2-Thinking-0905
Kimi K2-Thinking-0905: API Pricing, Context Window & Benchmarks
Kimi K2-Thinking-0905 is a language model from MoonshotAI, released in September 2025.
Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, it is built as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on
Kimi K2-Thinking-0905 benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Kimi K2-Thinking-0905 across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Kimi K2-Thinking-0905 holds up as conversations get longer.
Quality Tracker
Kimi K2-Thinking-0905 Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Kimi K2-Thinking-0905 pricing
Providers
Kimi K2-Thinking-0905 starts at $0.470 per million input tokens and $2.00 per million output tokens via DeepInfra. See all 3 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.470 | — | $2.00 | 262.1K/262.1K | — | — | / | |
| $0.480 | — | $2.00 | 262.1K/262.1K | — | — | / | |
| $0.600 | — | $2.50 | 262.1K/262.1K | — | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Kimi K2-Thinking-0905 model size
Kimi K2-Thinking-0905 has 1 trillion parameters and was trained on 15.5 trillion tokens. See how it compares to other models in the same parameter range.
Kimi K2-Thinking-0905 context window
Input and output token limits for Kimi K2-Thinking-0905, plus how it ranks on long-context understanding.
Kimi K2-Thinking-0905 API
Available from the model provider
Kimi K2-Thinking-0905 has an official provider API. It is not currently routed through the LLM Stats gateway.
Read the official API documentationKimi K2-Thinking-0905 latency
Kimi K2-Thinking-0905 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Kimi K2-Thinking-0905 examples
Recent arena outputs from Kimi K2-Thinking-0905, picked from the highest-ranked matchups.
Kimi K2-Thinking-0905 license
Kimi K2-Thinking-0905 is released under the MIT license, which permits commercial use, has 1.0T parameters.
- License
- MIT
- Commercial use allowed
- Parameters
- 1.0T
MIT License - allows commercial use
Kimi K2-Thinking-0905 resources
Official sources for Kimi K2-Thinking-0905: api documentation, official playground, paper or system card, official launch post, source repository, model weights.
Kimi K2-Thinking-0905 vs other models
The most-compared alternatives to Kimi K2-Thinking-0905 are Claude Opus 4.6, Gemini 3 Pro, Grok-3. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Kimi K2-Thinking-0905
Models ranked just above and below Kimi K2-Thinking-0905 by LLM Stats score.
FAQ
Common questions about Kimi K2-Thinking-0905.