The AI arena is free today

Open Superagent
MicrosoftReleased on Apr 30, 2025

Phi 4 Reasoning: API Pricing, Context Window & Benchmarks

Phi 4 Reasoning is a language model from Microsoft, released in April 2025.

Phi-4-reasoning is a state-of-the-art open-weight reasoning model finetuned from Phi-4 using supervised fine-tuning on a dataset of chain-of-thought traces and reinforcement learning. It focuses on math, science, and coding skills.

Phi 4 Reasoning benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Phi 4 Reasoning across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Phi 4 Reasoning holds up as conversations get longer.

Quality Tracker

Phi 4 Reasoning Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Aug 26 2026
Notice missing or incorrect data?

Phi 4 Reasoning model size

Phi 4 Reasoning has 14 billion parameters and was trained on 16 billion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
14B
16Btokens
1× tokens-to-params ratio
Medium (10–30B)
14B
1B7B70B405B

Phi 4 Reasoning API

Available from the model provider

Phi 4 Reasoning has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Phi 4 Reasoning latency

Phi 4 Reasoning time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Phi 4 Reasoning examples

Recent arena outputs from Phi 4 Reasoning, picked from the highest-ranked matchups.

Phi 4 Reasoning license

Phi 4 Reasoning is released under the MIT license, which permits commercial use, has 14.0B parameters, has a knowledge cutoff of March 2025.

License
MIT
Commercial use allowed
Parameters
14.0B
Knowledge cutoff
March 2025

MIT License - allows commercial use

Phi 4 Reasoning resources

Official sources for Phi 4 Reasoning: api documentation, paper or system card, official launch post, model weights.

Phi 4 Reasoning vs other models

The most-compared alternatives to Phi 4 Reasoning are Grok-2, DeepSeek R1 Distill Llama 70B, LongCat-Flash-Chat. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Phi 4 Reasoning

Models ranked just above and below Phi 4 Reasoning by LLM Stats score.

 

Grok-2

Score pending
 

DeepSeek R1 Distill Llama 70B

Score pending
 

LongCat-Flash-Chat

Score pending
 

Qwen3 30B A3B

Score pending
 

DeepSeek R1 Distill Qwen 14B

Score pending
 

LongCat-Flash-Lite

Score pending

FAQ

Common questions about Phi 4 Reasoning.

When was Phi 4 Reasoning released?

Phi 4 Reasoning was released on April 30, 2025 by Microsoft. This is the official Phi 4 Reasoning release date tracked on LLM Stats.

Is Phi 4 Reasoning available via API?

Yes, Phi 4 Reasoning is available via API. See the official documentation for authentication and endpoint details.

How big is Phi 4 Reasoning?

Phi 4 Reasoning has 14 billion parameters. It was trained on 16 billion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Phi 4 Reasoning?

Phi 4 Reasoning was created by Microsoft.

What is the license for Phi 4 Reasoning?

Phi 4 Reasoning is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Phi 4 Reasoning?

Phi 4 Reasoning has a knowledge cutoff of March 2025, meaning it was trained on data up to that point and may not know about events after it.

Where is the Phi 4 Reasoning paper or technical report?

Phi 4 Reasoning has a paper or technical report available at https://arxiv.org/abs/2504.21318. Use that source for architecture, training, release and evaluation details.

What models should I compare Phi 4 Reasoning against?

Common Phi 4 Reasoning comparisons include Phi 4 Reasoning vs Grok-2, Phi 4 Reasoning vs DeepSeek R1 Distill Llama 70B, Phi 4 Reasoning vs LongCat-Flash-Chat. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.