Kimi K3: API Pricing, Context Window & Benchmarks
Kimi K3 is a language model from MoonshotAI, released in July 2026, with multimodal input, a 1.0M-token context window, and pricing from $3.00/M input, $0.300/M cached input, $15.00/M output.
Kimi K3 is Moonshot AI's flagship open Mixture-of-Experts model for long-horizon coding, knowledge work, and reasoning. It has 2.8 trillion total parameters, activates 16 of 896 experts, and combines Kimi Delta Attention, Attention
Kimi K3 benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Kimi K3 across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Kimi K3 holds up as conversations get longer.
Quality Tracker
Kimi K3 Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Kimi K3 pricing
Providers
Kimi K3 starts at $3.00 per million input tokens and $15.00 per million output tokens via Fireworks. Reused prompt prefixes cost $0.300 per million cached input tokens. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $3.00 | $0.300 | $15.00 | 1.0M/1.0M | 0.00 | 3 | / | |
| $3.00 | $0.300 | $15.00 | 1.0M/1.0M | 7.41 | 57 | / | |
| $3.00 | $0.300 | $15.00 | 1.0M/1.0M | — | — | / | |
| $3.00 | $0.300 | $15.00 | 1.0M/1.0M | 1.92 | 10 | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Kimi K3 context window
Input and output token limits for Kimi K3, plus how it ranks on long-context understanding.
Kimi K3 API
Run a request to see the response
Use it in your code
Billed at $3.00 input / $0.30 cached input / $15.00 output per 1M tokens through the LLM Stats gateway.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://gateway.llm-stats.com/v1"
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "What is machine learning?"}
]
)
print(response.choices[0].message.content)Need an API key? Create one above in the playground, or read the API documentation.
Kimi K3 latency
Kimi K3 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Provider operational metrics
Time to first token, output throughput, and failed-request rate from live API traffic
Kimi K3 examples
Recent arena outputs from Kimi K3, picked from the highest-ranked matchups.
Kimi K3 license
Kimi K3 is a proprietary model available under its provider's product and API terms, has 2.8T parameters.
- License
- Kimi K3 License
- Hosted access
- Parameters
- 2.8T
Open-weights license permitting use, modification, distribution, sublicensing, sale, deployment, and fine-tuning. Model-as-a-Service businesses earning more than $20M in a consecutive 12-month period require a separate Moonshot AI agreement; very large commercial products must display Kimi K3 attribution.
Kimi K3 resources
Official sources for Kimi K3: api documentation, official playground, official launch post, source repository.
Kimi K3 vs other models
The most-compared alternatives to Kimi K3 are Gemini 3 Flash, GPT-5.2 Pro, Claude Opus 4.8. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Kimi K3
Models ranked just above and below Kimi K3 by LLM Stats score.
FAQ
Common questions about Kimi K3.