MicrosoftReleased on Feb 1, 2025

Phi-4-multimodal-instruct: API Pricing, Context Window & Benchmarks

Phi-4-multimodal-instruct is a language model from Microsoft, released in February 2025, with multimodal input.

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a

Input
TextImage
Output
Text

Phi-4-multimodal-instruct benchmarks

Rankings

Quality Tracker

Phi-4-multimodal-instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

Phi-4-multimodal-instruct pricing

Providers

Phi-4-multimodal-instruct starts at $0.0500 per million input tokens and $0.100 per million output tokens via DeepInfra.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.0500$0.100$0.0600128.0K/128.0K
/0.50
25/
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Phi-4-multimodal-instruct model size

Phi-4-multimodal-instruct has 5.6 billion parameters and was trained on 5 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
5.6B
5Ttokens
893× tokens-to-params ratio
Small (3–10B)
5.6B
1B7B70B405B

Phi-4-multimodal-instruct context window

Input and output token limits for Phi-4-multimodal-instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Phi-4-multimodal-instruct latency

Phi-4-multimodal-instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Phi-4-multimodal-instruct examples

Recent arena outputs from Phi-4-multimodal-instruct, picked from the highest-ranked matchups.

Phi-4-multimodal-instruct license

Phi-4-multimodal-instruct is released under the MIT license, which permits commercial use, has 5.6B parameters, has a knowledge cutoff of June 2024.

License
MIT
Commercial use allowed
Parameters
5.6B
Knowledge cutoff
June 2024

MIT License - allows commercial use

Phi-4-multimodal-instruct resources

Official sources for Phi-4-multimodal-instruct: official playground, paper or system card, official launch post, model weights.

Phi-4-multimodal-instruct vs other models

The most-compared alternatives to Phi-4-multimodal-instruct are GPT-4o, Grok-1.5V, Llama 3.2 90B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Phi-4-multimodal-instruct

Models ranked just above and below Phi-4-multimodal-instruct by LLM Stats score.

 

GPT-4o

Score pending
 

Grok-1.5V

Score pending
 

Llama 3.2 90B Instruct

Score pending
 

DeepSeek VL2

Score pending
 

Nova Lite

Score pending
 

DeepSeek VL2 Small

Score pending

FAQ

Common questions about Phi-4-multimodal-instruct.

When was Phi-4-multimodal-instruct released?

Phi-4-multimodal-instruct was released on February 1, 2025 by Microsoft. This is the official Phi-4-multimodal-instruct release date tracked on LLM Stats.

How much does Phi-4-multimodal-instruct cost?

Phi-4-multimodal-instruct pricing starts at $0.05 per million input tokens and $0.10 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct has 5.6 billion parameters. It was trained on 5.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct was created by Microsoft.

What is the license for Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct has a knowledge cutoff of June 2024, meaning it was trained on data up to that point and may not know about events after it.

Is Phi-4-multimodal-instruct multimodal?

Yes, Phi-4-multimodal-instruct is multimodal and can accept both text and images as input.

What is Phi-4-multimodal-instruct latency?

Phi-4-multimodal-instruct p95 time to first token is 0.50 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Phi-4-multimodal-instruct?

Phi-4-multimodal-instruct is available through 1 provider including DeepInfra.

Where is the Phi-4-multimodal-instruct paper or technical report?

Phi-4-multimodal-instruct has a paper or technical report available at https://arxiv.org/abs/2503.01743. Use that source for architecture, training, release and evaluation details.

What models should I compare Phi-4-multimodal-instruct against?

Common Phi-4-multimodal-instruct comparisons include Phi-4-multimodal-instruct vs GPT-4o, Phi-4-multimodal-instruct vs Grok-1.5V, Phi-4-multimodal-instruct vs Llama 3.2 90B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.