MetaReleased on Apr 5, 2025

Llama 4 Scout: API Pricing, Context Window & Benchmarks

Llama 4 Scout is a language model from Meta, released in April 2025, with multimodal input.

Llama 4 Scout is a natively multimodal model capable of processing both text and images. It features a 17 billion activated parameter (109B total) mixture-of-experts (MoE) architecture with 16 experts, supporting a wide range of multimodal

Input
TextImage
Output
Text

Llama 4 Scout benchmarks

Rankings

Quality Tracker

Llama 4 Scout Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Jul 20 2026
Notice missing or incorrect data?

Llama 4 Scout pricing

Providers

Llama 4 Scout starts at $0.0800 per million input tokens and $0.300 per million output tokens via DeepInfra. See all 6 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MWorkload 1M + 100KContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.0800$0.300$0.11010.0M/10.0M
/0.31
76/
/
Lambda logoLambda
$0.0800$0.300$0.11010.0M/10.0M
/0.43
140/
/
Novita logoNovita
$0.100$0.500$0.15010.0M/10.0M
/0.85
70/
/
Groq logoGroq
$0.110$0.340$0.14410.0M/10.0M
/1.08
776/
/
Fireworks logoFireworks
$0.150$0.600$0.21010.0M/10.0M
/0.53
116/
/
Together logoTogether
$0.180$0.590$0.23910.0M/10.0M
/0.54
107/
/

Workload cost uses 1M input tokens plus 100K output tokens at published list prices without assuming a cache hit. Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

Llama 4 Scout model size

Llama 4 Scout has 109 billion parameters and was trained on 40 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
109B
40Ttokens
367× tokens-to-params ratio
Very large (80–200B)
109B
1B7B70B405B

Llama 4 Scout context window

Input and output token limits for Llama 4 Scout, plus how it ranks on long-context understanding.

InputOutput
10Mtokens
10Mtokens
15.0k pages of text
10M
8K128K1M

Llama 4 Scout API

Available from the model provider

Llama 4 Scout has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Llama 4 Scout latency

Llama 4 Scout time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Llama 4 Scout examples

Recent arena outputs from Llama 4 Scout, picked from the highest-ranked matchups.

Llama 4 Scout license

Llama 4 Scout is released under the Llama 4 Community License Agreement license, which restricts commercial use, has 109.0B parameters.

License
Llama 4 Community License Agreement
Non-commercial
Parameters
109.0B

Llama 4 Scout resources

Official sources for Llama 4 Scout: api documentation, official playground, source repository.

Llama 4 Scout vs other models

The most-compared alternatives to Llama 4 Scout are Llama 3.1 405B Instruct, Grok-2, Phi 4 Reasoning. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 4 Scout

Models ranked just above and below Llama 4 Scout by LLM Stats score.

 

Llama 3.1 405B Instruct

Score pending
 

Grok-2

Score pending
 

Phi 4 Reasoning

Score pending
 

Ministral 3 (14B Base 2512)

Score pending
 

DeepSeek R1 Distill Qwen 14B

Score pending
 

Phi 4

Score pending

FAQ

Common questions about Llama 4 Scout.

When was Llama 4 Scout released?

Llama 4 Scout was released on April 5, 2025 by Meta. This is the official Llama 4 Scout release date tracked on LLM Stats.

How much does Llama 4 Scout cost?

Llama 4 Scout pricing starts at $0.08 per million input tokens and $0.30 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is Llama 4 Scout available via API?

Yes, Llama 4 Scout is available via API. See the official documentation for authentication and endpoint details. It is served by 6 providers tracked on LLM Stats.

How big is Llama 4 Scout?

Llama 4 Scout has 109 billion parameters. It was trained on 40.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 4 Scout?

Llama 4 Scout was created by Meta.

What is the license for Llama 4 Scout?

Llama 4 Scout is released under the Llama 4 Community License Agreement license. This is an open-source / open-weight license that permits self-hosting.

Is Llama 4 Scout multimodal?

Yes, Llama 4 Scout is multimodal and can accept both text and images as input.

What is Llama 4 Scout latency?

Llama 4 Scout p95 time to first token is 0.31 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Llama 4 Scout?

Llama 4 Scout is available through 6 providers including DeepInfra, Lambda, Novita, and 3 more.

What models should I compare Llama 4 Scout against?

Common Llama 4 Scout comparisons include Llama 4 Scout vs Llama 3.1 405B Instruct, Llama 4 Scout vs Grok-2, Llama 4 Scout vs Phi 4 Reasoning. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.