- Organizations
- Meta
- Llama 3.1 405B Instruct
Llama 3.1 405B Instruct: API Pricing, Context Window & Benchmarks
Llama 3.1 405B Instruct is a language model from Meta, released in July 2024.
Llama 3.1 405B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks. The model supports 8 languages and has a 128K
Llama 3.1 405B Instruct benchmarks
Rankings
Quality Tracker
Llama 3.1 405B Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Llama 3.1 405B Instruct pricing
Providers
Llama 3.1 405B Instruct starts at $0.890 per million input tokens and $0.890 per million output tokens via Lambda. See all 8 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Output $/M | Context in / out | TTFT p50 / p95 s | Output avg / p5 c/s | Success 7d | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.890 | $0.890 | 128.0K/128.0K | —/0.50 | 42/— | — | / | |
| $1.79 | $1.79 | 128.0K/128.0K | —/0.50 | 27/— | — | / | |
| $3.00 | $3.00 | 128.0K/128.0K | —/0.50 | 78/— | — | / | |
| $3.00 | $3.00 | 128.0K/128.0K | —/0.50 | 100/— | — | / | |
| $3.50 | $3.50 | 128.0K/128.0K | —/0.50 | 35/— | — | / | |
| $4.00 | $4.00 | 128.0K/128.0K | —/0.50 | 40/— | — | / | |
| $5.00 | $16.00 | 128.0K/128.0K | —/0.40 | 42/— | — | / | |
| $9.50 | $9.50 | 128.0K/128.0K | —/0.50 | 22/— | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.
Llama 3.1 405B Instruct model size
Llama 3.1 405B Instruct has 405 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.
Llama 3.1 405B Instruct context window
Input and output token limits for Llama 3.1 405B Instruct, plus how it ranks on long-context understanding.
Llama 3.1 405B Instruct API
Available from the model provider
Llama 3.1 405B Instruct has an official provider API. It is not currently routed through the LLM Stats gateway.
Read the official API documentationLlama 3.1 405B Instruct latency
Llama 3.1 405B Instruct time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.
Llama 3.1 405B Instruct examples
Recent arena outputs from Llama 3.1 405B Instruct, picked from the highest-ranked matchups.
Llama 3.1 405B Instruct license
Llama 3.1 405B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 405.0B parameters.
- License
- Llama 3.1 Community License
- Non-commercial
- Parameters
- 405.0B
Llama 3.1 405B Instruct resources
Official sources for Llama 3.1 405B Instruct: api documentation, official playground, official launch post, model weights.
Llama 3.1 405B Instruct vs other models
The most-compared alternatives to Llama 3.1 405B Instruct are Kimi-k1.5, Claude 3 Opus, Llama 3.3 70B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Llama 3.1 405B Instruct
Models ranked just above and below Llama 3.1 405B Instruct by LLM Stats score.
FAQ
Common questions about Llama 3.1 405B Instruct.