- Organizations
- MoonshotAI
- Kimi K2 Instruct
Kimi K2 Instruct: API Pricing, Context Window & Benchmarks
Kimi K2 Instruct is a language model from MoonshotAI, released in July 2025.
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the MuonClip optimizer, it achieves exceptional performance across frontier knowledge,
Kimi K2 Instruct benchmarks
Rankings
Quality Tracker
Kimi K2 Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Kimi K2 Instruct pricing
Providers
Kimi K2 Instruct starts at $0.500 per million input tokens and $0.500 per million output tokens via Fireworks. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Output $/M | Workload 1M + 100K | Context in / out | TTFT p50 / p95 s | Output avg / p5 c/s | Success 7d | Modalities in / out |
|---|---|---|---|---|---|---|---|---|
| $0.500 | $0.500 | $0.550 | 200.0K/200.0K | —/— | —/— | — | / | |
| $0.570 | $2.30 | $0.800 | 131.1K/131.1K | —/0.95 | 45/— | — | / |
Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.
Kimi K2 Instruct model size
Kimi K2 Instruct has 1 trillion parameters and was trained on 15.5 trillion tokens. See how it compares to other models in the same parameter range.
Kimi K2 Instruct context window
Input and output token limits for Kimi K2 Instruct, plus how it ranks on long-context understanding.
Kimi K2 Instruct API
Available from the model provider
Kimi K2 Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.
Read the official API documentationKimi K2 Instruct latency
Kimi K2 Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Kimi K2 Instruct examples
Recent arena outputs from Kimi K2 Instruct, picked from the highest-ranked matchups.
Kimi K2 Instruct license
Kimi K2 Instruct is released under the MIT license, which permits commercial use, has 1.0T parameters.
- License
- MIT
- Commercial use allowed
- Parameters
- 1.0T
MIT License - allows commercial use
Kimi K2 Instruct resources
Official sources for Kimi K2 Instruct: api documentation, official playground, official launch post, source repository, model weights.
Kimi K2 Instruct vs other models
The most-compared alternatives to Kimi K2 Instruct are MiMo-V2.5-Pro, LongCat-Flash-Chat, Qwen3 VL 235B A22B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Kimi K2 Instruct
Models ranked just above and below Kimi K2 Instruct by LLM Stats score.
FAQ
Common questions about Kimi K2 Instruct.