The AI arena is free today

Open Superagent

Kimi-k1.5 vs Phi-4-multimodal-instruct

Kimi-k1.5 leads the LLM Stats Score 17.2 to 2.8.

Moonshot AI · Microsoft · Updated for 2026

Which is better?

Kimi-k1.5 leads the overall LLM Stats Score 17.2 to 2.8, ranking #218 overall.

In the 2 individual benchmarks reported for both models, Kimi-k1.5 wins 2; this is a narrower head-to-head signal than the composite indexes.

Based on current LLM Stats indexes, shared benchmarks, pricing, and model metadata for 2026.

Choose Kimi-k1.5

  • overall performance matters — it scores 17.2 and ranks #218 on LLM Stats
  • your work emphasizes reasoning — it leads those capability indexes
  • you value its reported benchmark strengths — it wins 2 of 2 exact shared results

Choose Phi-4-multimodal-instruct

  • you want the most recent training data — it shipped Feb 2025
  • you need open weights you can self-host or fine-tune

At a glance

The differences that matter most.

Core performance indexes
17.2
#218
2.8
#312
16.4
#215
-0.2
#320
Cost, coverage & limits
Benchmark wins
2 of 2
0 of 2
Input price
— / M
$0.05 / M
Output price
— / M
$0.10 / M
Context window
128,000

Capability indexes

Additional strengths measured across groups of related public benchmarks

2 shared
Index
Kimi-k1.5
Phi-4-multimodal-instruct
11.0#111
1.6#172
14.7#92
3.6#141
Conservative TrueSkill rating · higher is betterHow scores work

Individual benchmarks

9 reported for Kimi-k1.5 · 15 for Phi-4-multimodal-instruct

2 shared

Kimi-k1.5 outperforms in 2 benchmarks (MathVista, MMMU), while Phi-4-multimodal-instruct is better at 0 benchmarks.

Kimi-k1.5 significantly outperforms across most benchmarks.

Mon Sep 21 2026 • llm-stats.com

Human preference

Blind head-to-head votes and playground preference scores

Context Window

Maximum input and output token capacity

Only Phi-4-multimodal-instruct specifies input context (128,000 tokens). Only Phi-4-multimodal-instruct specifies output context (128,000 tokens).

Moonshot AI
Kimi-k1.5
Input- tokens
Output- tokens
Microsoft
Phi-4-multimodal-instruct
Input128,000 tokens
Output128,000 tokens
Mon Sep 21 2026 • llm-stats.com

Input capabilities

Documented input modalities across available providers

Both Kimi-k1.5 and Phi-4-multimodal-instruct support multimodal inputs.

They are both capable of processing various types of data, offering versatility in application.

Kimi-k1.5

Text
Images
Audio
Video

Phi-4-multimodal-instruct

Text
Images
Audio
Video

License

Usage and distribution terms

Kimi-k1.5 is licensed under a proprietary license, while Phi-4-multimodal-instruct uses MIT.

License differences may affect how you can use these models in commercial or open-source projects.

Kimi-k1.5

Proprietary

Closed source

Phi-4-multimodal-instruct

MIT

Open weights

Release Timeline

When each model was launched

Kimi-k1.5 was released on 2025-01-20, while Phi-4-multimodal-instruct was released on 2025-02-01.

Phi-4-multimodal-instruct is 0 month newer than Kimi-k1.5.

Kimi-k1.5

Jan 20, 2025

1.7 years ago

Phi-4-multimodal-instruct

Feb 1, 2025

1.6 years ago

1w newer

Knowledge Cutoff

When training data ends

Phi-4-multimodal-instruct has a documented knowledge cutoff of 2024-06-01, while Kimi-k1.5's cutoff date is not specified.

We can confirm Phi-4-multimodal-instruct's training data extends to 2024-06-01, but cannot make a direct comparison without Kimi-k1.5's cutoff date.

Kimi-k1.5

Phi-4-multimodal-instruct

Jun 2024

Outputs Comparison

Notice missing or incorrect data?Start an Issue discussion

Judge for yourself.

Run your own prompts against Kimi-k1.5 and Phi-4-multimodal-instruct side-by-side, then vote on the output you prefer.

Kimi-k1.5
✓ Preferred
Phi-4-multimodal-instruct
Open in Playground

FAQ

Common questions about Kimi-k1.5 vs Phi-4-multimodal-instruct.

Which is better, Kimi-k1.5 or Phi-4-multimodal-instruct?

Kimi-k1.5 leads the LLM Stats Score 17.2 to 2.8. Kimi-k1.5 is made by Moonshot AI and Phi-4-multimodal-instruct is made by Microsoft. The best choice depends on your use case — compare their capability indexes, individual benchmarks, pricing, and limits above.

How does Kimi-k1.5 compare to Phi-4-multimodal-instruct in benchmarks?

Kimi-k1.5 scores MATH-500: 96.2%, CLUEWSC: 91.4%, C-Eval: 88.3%, MMLU: 87.4%, IFEval: 87.2%. Phi-4-multimodal-instruct scores ScienceQA Visual: 97.5%, DocVQA: 93.2%, MMBench: 86.7%, POPE: 85.6%, OCRBench: 84.4%.

What are the context window sizes for Kimi-k1.5 and Phi-4-multimodal-instruct?

Kimi-k1.5 supports an unknown number of tokens and Phi-4-multimodal-instruct supports 128K tokens. A larger context window lets you process longer documents, conversations, or codebases in a single request.

What are the main differences between Kimi-k1.5 and Phi-4-multimodal-instruct?

Key differences include LLM Stats Score (17.2 vs 2.8), licensing (Proprietary vs MIT). See the full comparison above for benchmark-by-benchmark results.

Who makes Kimi-k1.5 and Phi-4-multimodal-instruct?

Kimi-k1.5 is developed by Moonshot AI and Phi-4-multimodal-instruct is developed by Microsoft.