MetaReleased on Jul 23, 2024

Llama 3.1 405B Instruct: API Pricing, Context Window & Benchmarks

Llama 3.1 405B Instruct is a language model from Meta, released in July 2024.

Llama 3.1 405B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks. The model supports 8 languages and has a 128K

Input
Text
Output
Text

Llama 3.1 405B Instruct benchmarks

Rankings

Quality Tracker

Llama 3.1 405B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Aug 03 2026
Notice missing or incorrect data?

Llama 3.1 405B Instruct pricing

Providers

Llama 3.1 405B Instruct starts at $0.890 per million input tokens and $0.890 per million output tokens via Lambda. See all 8 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
Lambda logoLambda
$0.890$0.890128.0K/128.0K
/0.50
42/
/
DeepInfra logoDeepInfra
$1.79$1.79128.0K/128.0K
/0.50
27/
/
Fireworks logoFireworks
$3.00$3.00128.0K/128.0K
/0.50
78/
/
Bedrock logoBedrock
$3.00$3.00128.0K/128.0K
/0.50
100/
/
Together logoTogether
$3.50$3.50128.0K/128.0K
/0.50
35/
/
Hyperbolic logoHyperbolic
$4.00$4.00128.0K/128.0K
/0.50
40/
/
Google logoGoogle
$5.00$16.00128.0K/128.0K
/0.40
42/
/
Replicate logoReplicate
$9.50$9.50128.0K/128.0K
/0.50
22/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

Llama 3.1 405B Instruct model size

Llama 3.1 405B Instruct has 405 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
405B
15Ttokens
37× tokens-to-params ratio
Frontier (200B+)
405B
1B7B70B405B

Llama 3.1 405B Instruct context window

Input and output token limits for Llama 3.1 405B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Llama 3.1 405B Instruct API

Available from the model provider

Llama 3.1 405B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

Llama 3.1 405B Instruct latency

Llama 3.1 405B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

Llama 3.1 405B Instruct examples

Recent arena outputs from Llama 3.1 405B Instruct, picked from the highest-ranked matchups.

Llama 3.1 405B Instruct license

Llama 3.1 405B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 405.0B parameters.

License
Llama 3.1 Community License
Non-commercial
Parameters
405.0B

Llama 3.1 405B Instruct resources

Official sources for Llama 3.1 405B Instruct: api documentation, official playground, official launch post, model weights.

Llama 3.1 405B Instruct vs other models

The most-compared alternatives to Llama 3.1 405B Instruct are Kimi-k1.5, Claude 3 Opus, Llama 3.3 70B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 3.1 405B Instruct

Models ranked just above and below Llama 3.1 405B Instruct by LLM Stats score.

 

Kimi-k1.5

Score pending
 

Claude 3 Opus

Score pending
 

Llama 3.3 70B Instruct

Score pending
 

GPT-4o

Score pending
 

Phi 4 Reasoning

Score pending
 

Grok-2 mini

Score pending

FAQ

Common questions about Llama 3.1 405B Instruct.

When was Llama 3.1 405B Instruct released?

Llama 3.1 405B Instruct was released on July 23, 2024 by Meta. This is the official Llama 3.1 405B Instruct release date tracked on LLM Stats.

How much does Llama 3.1 405B Instruct cost?

Llama 3.1 405B Instruct pricing starts at $0.89 per million input tokens and $0.89 per million output tokens via Lambda, the lowest price among tracked providers.

Is Llama 3.1 405B Instruct available via API?

Yes, Llama 3.1 405B Instruct is available via API. See the official documentation for authentication and endpoint details. It is served by 8 providers tracked on LLM Stats.

How big is Llama 3.1 405B Instruct?

Llama 3.1 405B Instruct has 405 billion parameters. It was trained on 15.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 3.1 405B Instruct?

Llama 3.1 405B Instruct was created by Meta.

What is the license for Llama 3.1 405B Instruct?

Llama 3.1 405B Instruct is released under the Llama 3.1 Community License license. This is an open-source / open-weight license that permits self-hosting.

What is Llama 3.1 405B Instruct latency?

Llama 3.1 405B Instruct p95 time to first token is 0.40 seconds via Google over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use Llama 3.1 405B Instruct?

Llama 3.1 405B Instruct is available through 8 providers including Lambda, DeepInfra, Fireworks, and 5 more.

Where is the Llama 3.1 405B Instruct paper or technical report?

Llama 3.1 405B Instruct has a paper or technical report available at https://ai.meta.com/blog/meta-llama-3-1/. Use that source for architecture, training, release and evaluation details.

What models should I compare Llama 3.1 405B Instruct against?

Common Llama 3.1 405B Instruct comparisons include Llama 3.1 405B Instruct vs Kimi-k1.5, Llama 3.1 405B Instruct vs Claude 3 Opus, Llama 3.1 405B Instruct vs Llama 3.3 70B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.