- Organizations
- Meta
- Llama 3.1 70B Instruct
Llama 3.1 70B Instruct: Benchmarks, Pricing & Context Window
Llama 3.1 70B Instruct is a language model from Meta, released in July 2024.
Llama 3.1 70B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks.
Llama 3.1 70B Instruct benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Llama 3.1 70B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Llama 3.1 70B Instruct holds up as conversations get longer.
Quality Tracker
Llama 3.1 70B Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Llama 3.1 70B Instruct pricing
Providers
Llama 3.1 70B Instruct starts at $0.200 per million input tokens and $0.200 per million output tokens via Lambda. See all 9 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.200 | — | $0.200 | 128.0K/128.0K | 0.50 | — | / | |
| $0.350 | — | $0.400 | 128.0K/128.0K | 0.50 | — | / | |
| $0.400 | — | $0.400 | 128.0K/128.0K | 0.50 | — | / | |
| $0.590 | — | $0.780 | 128.0K/128.0K | 0.50 | — | / | |
| $0.600 | — | $0.600 | 128.0K/128.0K | 0.20 | — | / | |
| $0.890 | — | $0.890 | 128.0K/128.0K | 0.50 | — | / | |
| $0.890 | — | $0.890 | 128.0K/128.0K | 0.50 | — | / | |
| $0.890 | — | $0.890 | 128.0K/128.0K | 0.50 | — | / | |
| $5.00 | — | $10.00 | 128.0K/128.0K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Llama 3.1 70B Instruct model size
Llama 3.1 70B Instruct has 70 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.
Llama 3.1 70B Instruct context window
Input and output token limits for Llama 3.1 70B Instruct, plus how it ranks on long-context understanding.
Try now
Make it with
Llama 3.1 70B Instruct.
Llama 3.1 70B Instruct latency
Llama 3.1 70B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Llama 3.1 70B Instruct examples
Recent arena outputs from Llama 3.1 70B Instruct, picked from the highest-ranked matchups.
Llama 3.1 70B Instruct license
Llama 3.1 70B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 70.0B parameters.
- License
- Llama 3.1 Community License
- Non-commercial
- Parameters
- 70.0B
Llama 3.1 70B Instruct resources
Official sources for Llama 3.1 70B Instruct: provider documentation, paper or system card, official launch post, source repository, model weights.
Llama 3.1 70B Instruct vs other models
The most-compared alternatives to Llama 3.1 70B Instruct are Qwen3 VL 235B A22B Instruct, Nova Lite, Mistral Large 2. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Llama 3.1 70B Instruct
Models ranked just above and below Llama 3.1 70B Instruct by LLM Stats score.
FAQ
Common questions about Llama 3.1 70B Instruct.