MicrosoftReleased on Aug 23, 2024

Phi-3.5-vision-instruct: API Pricing, Context Window & Benchmarks

Phi-3.5-vision-instruct is a language model from Microsoft, released in August 2024, with multimodal input.

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image

Phi-3.5-vision-instruct benchmarks

Rankings

Quality Tracker

Phi-3.5-vision-instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 27 2026
Notice missing or incorrect data?

Phi-3.5-vision-instruct model size

Phi-3.5-vision-instruct has 4.2 billion parameters and was trained on 500 billion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
4.2B
500Btokens
119× tokens-to-params ratio
Small (3–10B)
4.2B
1B7B70B405B

Phi-3.5-vision-instruct API

Available from the model provider

Phi-3.5-vision-instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Phi-3.5-vision-instruct latency

Phi-3.5-vision-instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Phi-3.5-vision-instruct examples

Recent arena outputs from Phi-3.5-vision-instruct, picked from the highest-ranked matchups.

Phi-3.5-vision-instruct license

Phi-3.5-vision-instruct is released under the MIT license, which permits commercial use, has 4.2B parameters.

License
MIT
Commercial use allowed
Parameters
4.2B

MIT License - allows commercial use

Phi-3.5-vision-instruct resources

Official sources for Phi-3.5-vision-instruct: api documentation, paper or system card, official launch post.

Phi-3.5-vision-instruct vs other models

The most-compared alternatives to Phi-3.5-vision-instruct are Phi-4-multimodal-instruct, DeepSeek VL2, DeepSeek VL2 Small. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Phi-3.5-vision-instruct

Models ranked just above and below Phi-3.5-vision-instruct by LLM Stats score.

 

Phi-4-multimodal-instruct

Score pending
 

DeepSeek VL2

Score pending
 

DeepSeek VL2 Small

Score pending
 

Llama 3.2 11B Instruct

Score pending
 

DeepSeek VL2 Tiny

Score pending
 

Gemini 1.0 Pro

Score pending

FAQ

Common questions about Phi-3.5-vision-instruct.

When was Phi-3.5-vision-instruct released?

Phi-3.5-vision-instruct was released on August 23, 2024 by Microsoft. This is the official Phi-3.5-vision-instruct release date tracked on LLM Stats.

Is Phi-3.5-vision-instruct available via API?

Yes, Phi-3.5-vision-instruct is available via API. See the official documentation for authentication and endpoint details.

How big is Phi-3.5-vision-instruct?

Phi-3.5-vision-instruct has 4.2 billion parameters. It was trained on 500 billion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Phi-3.5-vision-instruct?

Phi-3.5-vision-instruct was created by Microsoft.

What is the license for Phi-3.5-vision-instruct?

Phi-3.5-vision-instruct is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

Is Phi-3.5-vision-instruct multimodal?

Yes, Phi-3.5-vision-instruct is multimodal and can accept both text and images as input.

Where is the Phi-3.5-vision-instruct paper or technical report?

Phi-3.5-vision-instruct has a paper or technical report available at https://arxiv.org/abs/2404.14219. Use that source for architecture, training, release and evaluation details.

What models should I compare Phi-3.5-vision-instruct against?

Common Phi-3.5-vision-instruct comparisons include Phi-3.5-vision-instruct vs Phi-4-multimodal-instruct, Phi-3.5-vision-instruct vs DeepSeek VL2, Phi-3.5-vision-instruct vs DeepSeek VL2 Small. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.