- Organizations
- Microsoft
- Phi-4-multimodal-instruct
Phi-4-multimodal-instruct: API Pricing, Context Window & Benchmarks
Phi-4-multimodal-instruct is a language model from Microsoft, released in February 2025, with multimodal input.
Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a
Phi-4-multimodal-instruct benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Phi-4-multimodal-instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Phi-4-multimodal-instruct holds up as conversations get longer.
Quality Tracker
Phi-4-multimodal-instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Phi-4-multimodal-instruct pricing
Providers
Phi-4-multimodal-instruct starts at $0.0500 per million input tokens and $0.100 per million output tokens via DeepInfra.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0500 | — | $0.100 | 128.0K/128.0K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Phi-4-multimodal-instruct model size
Phi-4-multimodal-instruct has 5.6 billion parameters and was trained on 5 trillion tokens. See how it compares to other models in the same parameter range.
Phi-4-multimodal-instruct context window
Input and output token limits for Phi-4-multimodal-instruct, plus how it ranks on long-context understanding.
Phi-4-multimodal-instruct latency
Phi-4-multimodal-instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Phi-4-multimodal-instruct examples
Recent arena outputs from Phi-4-multimodal-instruct, picked from the highest-ranked matchups.
Phi-4-multimodal-instruct license
Phi-4-multimodal-instruct is released under the MIT license, which permits commercial use, has 5.6B parameters, has a knowledge cutoff of June 2024.
- License
- MIT
- Commercial use allowed
- Parameters
- 5.6B
- Knowledge cutoff
- June 2024
MIT License - allows commercial use
Phi-4-multimodal-instruct resources
Official sources for Phi-4-multimodal-instruct: official playground, paper or system card, official launch post, model weights.
Phi-4-multimodal-instruct vs other models
The most-compared alternatives to Phi-4-multimodal-instruct are GPT-4o, Llama 3.2 90B Instruct, DeepSeek VL2. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Phi-4-multimodal-instruct
Models ranked just above and below Phi-4-multimodal-instruct by LLM Stats score.
FAQ
Common questions about Phi-4-multimodal-instruct.