The AI arena is free today

Open Superagent
MetaReleased on Apr 5, 2025

Llama 4 Maverick: API Pricing, Context Window & Benchmarks

Llama 4 Maverick is a language model from Meta, released in April 2025, with multimodal input.

Llama 4 Maverick is a natively multimodal model capable of processing both text and images. It features a 17 billion active parameter mixture-of-experts (MoE) architecture with 128 experts, supporting a wide range of multimodal tasks such

Input
TextImage
Output
Text

Llama 4 Maverick benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Llama 4 Maverick across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Llama 4 Maverick holds up as conversations get longer.

Quality Tracker

Llama 4 Maverick Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Aug 31 2026
Notice missing or incorrect data?

Llama 4 Maverick pricing

Providers

Llama 4 Maverick starts at $0.170 per million input tokens and $0.600 per million output tokens via DeepInfra. See all 7 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.170$0.6001.0M/1.0M
0.38
/
Novita logoNovita
$0.170$0.8501.0M/1.0M
0.62
/
Lambda logoLambda
$0.180$0.6001.0M/1.0M
0.65
/
Groq logoGroq
$0.200$0.6001.0M/1.0M
0.27
/
Fireworks logoFireworks
$0.220$0.8801.0M/1.0M
0.62
/
Together logoTogether
$0.270$0.8501.0M/1.0M
0.20
/
Sambanova logoSambanova
$0.630$1.791.0M/1.0M
2.04
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

Llama 4 Maverick model size

Llama 4 Maverick has 400 billion parameters and was trained on 22 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
400B
22Ttokens
55× tokens-to-params ratio
Frontier (200B+)
400B
1B7B70B405B

Llama 4 Maverick context window

Input and output token limits for Llama 4 Maverick, plus how it ranks on long-context understanding.

InputOutput
1Mtokens
1Mtokens
1.5k pages of text
1M
8K128K1M

Llama 4 Maverick API

Available from the model provider

Llama 4 Maverick has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Llama 4 Maverick latency

Llama 4 Maverick time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Llama 4 Maverick examples

Recent arena outputs from Llama 4 Maverick, picked from the highest-ranked matchups.

Llama 4 Maverick license

Llama 4 Maverick is released under the Llama 4 Community License Agreement license, which restricts commercial use, has 400.0B parameters.

License
Llama 4 Community License Agreement
Non-commercial
Parameters
400.0B

Llama 4 Maverick resources

Official sources for Llama 4 Maverick: api documentation, official playground, source repository.

Llama 4 Maverick vs other models

The most-compared alternatives to Llama 4 Maverick are LongCat-Flash-Chat, o1-mini, DeepSeek-V3 0324. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 4 Maverick

Models ranked just above and below Llama 4 Maverick by LLM Stats score.

 

LongCat-Flash-Chat

Score pending
 

o1-mini

Score pending
 

DeepSeek-V3 0324

Score pending
 

Qwen3 VL 30B A3B Thinking

Score pending
 

Qwen3 VL 8B Thinking

Score pending
 

LongCat-Flash-Lite

Score pending

FAQ

Common questions about Llama 4 Maverick.

When was Llama 4 Maverick released?

Llama 4 Maverick was released on April 5, 2025 by Meta. This is the official Llama 4 Maverick release date tracked on LLM Stats.

How much does Llama 4 Maverick cost?

Llama 4 Maverick pricing starts at $0.17 per million input tokens and $0.60 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is Llama 4 Maverick available via API?

Yes, Llama 4 Maverick is available via API. See the official documentation for authentication and endpoint details. It is served by 7 providers tracked on LLM Stats.

How big is Llama 4 Maverick?

Llama 4 Maverick has 400 billion parameters. It was trained on 22.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 4 Maverick?

Llama 4 Maverick was created by Meta.

What is the license for Llama 4 Maverick?

Llama 4 Maverick is released under the Llama 4 Community License Agreement license. This is an open-source / open-weight license that permits self-hosting.

Is Llama 4 Maverick multimodal?

Yes, Llama 4 Maverick is multimodal and can accept both text and images as input.

What is Llama 4 Maverick latency?

Llama 4 Maverick p95 time to first token is 0.20 seconds via Together over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Llama 4 Maverick?

Llama 4 Maverick is available through 7 providers including DeepInfra, Novita, Lambda, and 4 more.

What models should I compare Llama 4 Maverick against?

Common Llama 4 Maverick comparisons include Llama 4 Maverick vs LongCat-Flash-Chat, Llama 4 Maverick vs o1-mini, Llama 4 Maverick vs DeepSeek-V3 0324. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.