MoonshotAIReleased on Jul 11, 2025

Kimi K2 Instruct: API Pricing, Context Window & Benchmarks

Kimi K2 Instruct is a language model from MoonshotAI, released in July 2025.

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the MuonClip optimizer, it achieves exceptional performance across frontier knowledge,

Input
Text
Output
Text

Kimi K2 Instruct benchmarks

Rankings

Quality Tracker

Kimi K2 Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

Kimi K2 Instruct pricing

Providers

Kimi K2 Instruct starts at $0.500 per million input tokens and $0.500 per million output tokens via Fireworks. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Fireworks logoFireworks
$0.500$0.500$0.550200.0K/200.0K
/
/
/
Novita logoNovita
$0.570$2.30$0.800131.1K/131.1K
/0.95
45/
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

Kimi K2 Instruct model size

Kimi K2 Instruct has 1 trillion parameters and was trained on 15.5 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
1T
15.5Ttokens
16× tokens-to-params ratio
Frontier (200B+)
1T
1B7B70B405B

Kimi K2 Instruct context window

Input and output token limits for Kimi K2 Instruct, plus how it ranks on long-context understanding.

InputOutput
200Ktokens
200Ktokens
301 pages of text
200K
8K128K1M

Kimi K2 Instruct API

Available from the model provider

Kimi K2 Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Kimi K2 Instruct latency

Kimi K2 Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Kimi K2 Instruct examples

Recent arena outputs from Kimi K2 Instruct, picked from the highest-ranked matchups.

Kimi K2 Instruct license

Kimi K2 Instruct is released under the MIT license, which permits commercial use, has 1.0T parameters.

License
MIT
Commercial use allowed
Parameters
1.0T

MIT License - allows commercial use

Kimi K2 Instruct resources

Official sources for Kimi K2 Instruct: api documentation, official playground, official launch post, source repository, model weights.

Kimi K2 Instruct vs other models

The most-compared alternatives to Kimi K2 Instruct are MiMo-V2.5-Pro, LongCat-Flash-Chat, Qwen3 VL 235B A22B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Kimi K2 Instruct

Models ranked just above and below Kimi K2 Instruct by LLM Stats score.

 

MiMo-V2.5-Pro

Score pending
 

LongCat-Flash-Chat

Score pending
 

Qwen3 VL 235B A22B Instruct

Score pending
 

DeepSeek-V3 0324

Score pending
 

Claude Sonnet 4

Score pending
 

Kimi K2-Instruct-0905

Score pending

FAQ

Common questions about Kimi K2 Instruct.

When was Kimi K2 Instruct released?

Kimi K2 Instruct was released on July 11, 2025 by MoonshotAI. This is the official Kimi K2 Instruct release date tracked on LLM Stats.

How much does Kimi K2 Instruct cost?

Kimi K2 Instruct pricing starts at $0.50 per million input tokens and $0.50 per million output tokens via Fireworks, the lowest price among tracked providers.

Is Kimi K2 Instruct available via API?

Yes, Kimi K2 Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 2 providers tracked on LLM Stats.

How big is Kimi K2 Instruct?

Kimi K2 Instruct has 1000 billion parameters. It was trained on 15.5 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Kimi K2 Instruct?

Kimi K2 Instruct was created by MoonshotAI.

What is the license for Kimi K2 Instruct?

Kimi K2 Instruct is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is Kimi K2 Instruct latency?

Kimi K2 Instruct p95 time to first token is 0.95 seconds via Novita over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Kimi K2 Instruct?

Kimi K2 Instruct is available through 2 providers including Fireworks, Novita.

Where is the Kimi K2 Instruct paper or technical report?

Kimi K2 Instruct has a paper or technical report available at https://moonshotai.github.io/Kimi-K2/. Use that source for architecture, training, release and evaluation details.

What models should I compare Kimi K2 Instruct against?

Common Kimi K2 Instruct comparisons include Kimi K2 Instruct vs MiMo-V2.5-Pro, Kimi K2 Instruct vs LongCat-Flash-Chat, Kimi K2 Instruct vs Qwen3 VL 235B A22B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.