MicrosoftReleased on Dec 12, 2024

Phi 4: API Pricing, Context Window & Benchmarks

Phi 4 is a language model from Microsoft, released in December 2024.

phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and

Input
Text
Output
Text

Phi 4 benchmarks

Rankings

Quality Tracker

Phi 4 Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

Phi 4 pricing

Providers

Phi 4 starts at $0.0700 per million input tokens and $0.140 per million output tokens via DeepInfra.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.0700$0.140$0.084016.0K/16.0K
/0.20
33/
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Phi 4 model size

Phi 4 has 14.7 billion parameters and was trained on 9.8 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
14.7B
9.8Ttokens
667× tokens-to-params ratio
Medium (10–30B)
14.7B
1B7B70B405B

Phi 4 context window

Input and output token limits for Phi 4, plus how it ranks on long-context understanding.

InputOutput
16Ktokens
16Ktokens
24 pages of text
16K
8K128K1M

Phi 4 API

Available from the model provider

Phi 4 has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Phi 4 latency

Phi 4 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Phi 4 examples

Recent arena outputs from Phi 4, picked from the highest-ranked matchups.

Phi 4 license

Phi 4 is released under the MIT license, which permits commercial use, has 14.7B parameters, has a knowledge cutoff of June 2024.

License
MIT
Commercial use allowed
Parameters
14.7B
Knowledge cutoff
June 2024

MIT License - allows commercial use

Phi 4 resources

Official sources for Phi 4: api documentation, paper or system card, official launch post.

Phi 4 vs other models

The most-compared alternatives to Phi 4 are Grok-2, Llama 3.1 70B Instruct, Claude 3.5 Sonnet. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Phi 4

Models ranked just above and below Phi 4 by LLM Stats score.

 

Grok-2

Score pending
 

Llama 3.1 70B Instruct

Score pending
 

Claude 3.5 Sonnet

Score pending
 

Mistral Large 2

Score pending
 

Qwen3 VL 30B A3B Instruct

Score pending
 

Sarvam-30B

Score pending

FAQ

Common questions about Phi 4.

When was Phi 4 released?

Phi 4 was released on December 12, 2024 by Microsoft. This is the official Phi 4 release date tracked on LLM Stats.

How much does Phi 4 cost?

Phi 4 pricing starts at $0.07 per million input tokens and $0.14 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is Phi 4 available via API?

Yes, Phi 4 is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is Phi 4?

Phi 4 has 14.7 billion parameters. It was trained on 9.8 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Phi 4?

Phi 4 was created by Microsoft.

What is the license for Phi 4?

Phi 4 is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Phi 4?

Phi 4 has a knowledge cutoff of June 2024, meaning it was trained on data up to that point and may not know about events after it.

What is Phi 4 latency?

Phi 4 p95 time to first token is 0.20 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Phi 4?

Phi 4 is available through 1 provider including DeepInfra.

Where is the Phi 4 paper or technical report?

Phi 4 has a paper or technical report available at https://arxiv.org/pdf/2412.08905. Use that source for architecture, training, release and evaluation details.

What models should I compare Phi 4 against?

Common Phi 4 comparisons include Phi 4 vs Grok-2, Phi 4 vs Llama 3.1 70B Instruct, Phi 4 vs Claude 3.5 Sonnet. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.