XiaomiReleased on Apr 27, 2026

MiMo-V2.5-Pro: API Pricing, Context Window & Benchmarks

MiMo-V2.5-Pro is a language model from Xiaomi, released in April 2026, with a 1.0M-token context window, and pricing from $1.00/M input and $3.00/M output.

MiMo-V2.5-Pro is Xiaomi's 1.02T-parameter sparse Mixture-of-Experts language model with 42B active parameters and a 1M-token context window. It inherits the MiMo-V2-Flash hybrid-attention and Multi-Token Prediction design, extends context

Input
Text
Output
Text

MiMo-V2.5-Pro benchmarks

Rankings

Quality Tracker

MiMo-V2.5-Pro Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sat Jul 18 2026
Notice missing or incorrect data?

MiMo-V2.5-Pro pricing

Providers

MiMo-V2.5-Pro starts at $0.435 per million input tokens and $0.870 per million output tokens via Xiaomi. See all 3 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Xiaomi logoXiaomi
$0.435$0.870$0.5221.0M/131.1K
2.82/10.74
132/5
94.59%(37)
/
DeepInfra logoDeepInfra
$1.00$3.00$1.301.0M/131.1K
0.62/8.32
192/34
98.68%(76)
/
Novita logoNovita
$2.00$6.00$2.601.0M/131.1K
1.21/1.77
285/249
14.29%(21)
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

MiMo-V2.5-Pro model size

MiMo-V2.5-Pro has 1.0 trillion parameters and was trained on 27 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
1.0TMoE
27Ttokens
26× tokens-to-params ratio
Frontier (200B+)
1.0T
1B7B70B405B

MiMo-V2.5-Pro context window

Input and output token limits for MiMo-V2.5-Pro, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
131Ktokens
1.6k pages of text
1.0M
8K128K1M

MiMo-V2.5-Pro API

POST/v1/chat/completions

Run a request to see the response

Use it in your code

Billed at $1.00 input / $3.00 output per 1M tokens through the LLM Stats gateway.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://gateway.llm-stats.com/v1"
)

response = client.chat.completions.create(
    model="mimo-v2.5-pro",
    messages=[
        {"role": "user", "content": "What is machine learning?"}
    ]
)

print(response.choices[0].message.content)

Need an API key? Create one above in the playground, or read the API documentation.

MiMo-V2.5-Pro latency

MiMo-V2.5-Pro time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Provider operational metrics

Time to first token, output throughput, and failed-request rate from live API traffic

Loading chart...
Loading chart...
Loading chart...

MiMo-V2.5-Pro examples

Recent arena outputs from MiMo-V2.5-Pro, picked from the highest-ranked matchups.

MiMo-V2.5-Pro license

MiMo-V2.5-Pro is released under the MIT license, which permits commercial use, has 1.0T parameters.

License
MIT
Commercial use allowed
Parameters
1.0T

MIT License - allows commercial use

MiMo-V2.5-Pro resources

Official sources for MiMo-V2.5-Pro: api documentation, official playground, model weights.

MiMo-V2.5-Pro vs other models

The most-compared alternatives to MiMo-V2.5-Pro are Claude 3 Opus, GPT-4o, Qwen3 VL 235B A22B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like MiMo-V2.5-Pro

Models ranked just above and below MiMo-V2.5-Pro by LLM Stats score.

 

Claude 3 Opus

Score pending
 

GPT-4o

Score pending
 

Qwen3 VL 235B A22B Instruct

Score pending
 

Ministral 3 (8B Reasoning 2512)

Score pending
 

Qwen3.6 Plus

Score pending
 

LongCat-Flash-Lite

Score pending

FAQ

Common questions about MiMo-V2.5-Pro.

When was MiMo-V2.5-Pro released?

MiMo-V2.5-Pro was released on April 27, 2026 by Xiaomi. This is the official MiMo-V2.5-Pro release date tracked on LLM Stats.

How much does MiMo-V2.5-Pro cost?

MiMo-V2.5-Pro costs $1.00 per million input tokens and $3.00 per million output tokens through the LLM Stats API, which works with any OpenAI-compatible SDK. Across tracked providers, the lowest price is $0.43 per million input tokens via Xiaomi.

Is MiMo-V2.5-Pro available via API?

Yes. MiMo-V2.5-Pro is available through the LLM Stats API and works with any OpenAI-compatible SDK — point your client at the gateway base URL and pass the model name. It is served by 3 providers tracked on LLM Stats.

How big is MiMo-V2.5-Pro?

MiMo-V2.5-Pro has 1023.2 billion parameters. It was trained on 27.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created MiMo-V2.5-Pro?

MiMo-V2.5-Pro was created by Xiaomi.

What is the license for MiMo-V2.5-Pro?

MiMo-V2.5-Pro is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is MiMo-V2.5-Pro latency?

MiMo-V2.5-Pro p95 time to first token is 1.77 seconds via Novita over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use MiMo-V2.5-Pro?

MiMo-V2.5-Pro is available through 3 providers including Xiaomi, DeepInfra, Novita.

Where is the MiMo-V2.5-Pro paper or technical report?

MiMo-V2.5-Pro has a paper or technical report available at https://mimo.xiaomi.com/mimo-v2-5-pro/. Use that source for architecture, training, release and evaluation details.

What models should I compare MiMo-V2.5-Pro against?

Common MiMo-V2.5-Pro comparisons include MiMo-V2.5-Pro vs Claude 3 Opus, MiMo-V2.5-Pro vs GPT-4o, MiMo-V2.5-Pro vs Qwen3 VL 235B A22B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.