The AI arena is free today

Open Superagent
MicrosoftReleased on Feb 1, 2025

Phi-4-multimodal-instruct: API Pricing, Context Window & Benchmarks

Phi-4-multimodal-instruct is a language model from Microsoft, released in February 2025, with multimodal input.

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a

Input
TextImage
Output
Text

Phi-4-multimodal-instruct benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Phi-4-multimodal-instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Phi-4-multimodal-instruct holds up as conversations get longer.

Quality Tracker

Phi-4-multimodal-instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Sep 03 2026
Notice missing or incorrect data?

Phi-4-multimodal-instruct pricing

Providers

Phi-4-multimodal-instruct starts at $0.0500 per million input tokens and $0.100 per million output tokens via DeepInfra.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0500$0.100128.0K/128.0K
0.50
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Phi-4-multimodal-instruct model size

Phi-4-multimodal-instruct has 5.6 billion parameters and was trained on 5 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
5.6B
5Ttokens
893× tokens-to-params ratio
Small (3–10B)
5.6B
1B7B70B405B

Phi-4-multimodal-instruct context window

Input and output token limits for Phi-4-multimodal-instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Phi-4-multimodal-instruct latency

Phi-4-multimodal-instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Phi-4-multimodal-instruct examples

Recent arena outputs from Phi-4-multimodal-instruct, picked from the highest-ranked matchups.

Phi-4-multimodal-instruct license

Phi-4-multimodal-instruct is released under the MIT license, which permits commercial use, has 5.6B parameters, has a knowledge cutoff of June 2024.

License
MIT
Commercial use allowed
Parameters
5.6B
Knowledge cutoff
June 2024

MIT License - allows commercial use

Phi-4-multimodal-instruct resources

Official sources for Phi-4-multimodal-instruct: official playground, paper or system card, official launch post, model weights.

Phi-4-multimodal-instruct vs other models

The most-compared alternatives to Phi-4-multimodal-instruct are GPT-4o, Llama 3.2 90B Instruct, DeepSeek VL2. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Phi-4-multimodal-instruct

Models ranked just above and below Phi-4-multimodal-instruct by LLM Stats score.

 

GPT-4o

Score pending
 

Llama 3.2 90B Instruct

Score pending
 

DeepSeek VL2

Score pending
 

Nova Lite

Score pending
 

DeepSeek VL2 Small

Score pending
 

Qwen2.5 VL 7B Instruct

Score pending

FAQ

Common questions about Phi-4-multimodal-instruct.

When was Phi-4-multimodal-instruct released?

Phi-4-multimodal-instruct was released on February 1, 2025 by Microsoft. This is the official Phi-4-multimodal-instruct release date tracked on LLM Stats.

How much does Phi-4-multimodal-instruct cost?

Phi-4-multimodal-instruct pricing starts at $0.05 per million input tokens and $0.10 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct has 5.6 billion parameters. It was trained on 5.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct was created by Microsoft.

What is the license for Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct has a knowledge cutoff of June 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Phi-4-multimodal-instruct multimodal?

Yes, Phi-4-multimodal-instruct is multimodal and can accept both text and images as input.

What is Phi-4-multimodal-instruct latency?

Phi-4-multimodal-instruct p95 time to first token is 0.50 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct is available through 1 provider including DeepInfra.

Where is the Phi-4-multimodal-instruct paper or technical report?

Phi-4-multimodal-instruct has a paper or technical report available at https://arxiv.org/abs/2503.01743. Use that source for architecture, training, release and evaluation details.

What models should I compare Phi-4-multimodal-instruct against?

Common Phi-4-multimodal-instruct comparisons include Phi-4-multimodal-instruct vs GPT-4o, Phi-4-multimodal-instruct vs Llama 3.2 90B Instruct, Phi-4-multimodal-instruct vs DeepSeek VL2. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.