The AI arena is free today

Open Superagent
MetaReleased on Jul 23, 2024

Llama 3.1 70B Instruct: Benchmarks, Pricing & Context Window

Llama 3.1 70B Instruct is a language model from Meta, released in July 2024.

Llama 3.1 70B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks.

Input
Text
Output
Text

Llama 3.1 70B Instruct benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Llama 3.1 70B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Llama 3.1 70B Instruct holds up as conversations get longer.

Quality Tracker

Llama 3.1 70B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Wed Sep 16 2026
Notice missing or incorrect data?

Llama 3.1 70B Instruct pricing

Providers

Llama 3.1 70B Instruct starts at $0.200 per million input tokens and $0.200 per million output tokens via Lambda. See all 9 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
Lambda logoLambda
$0.200$0.200128.0K/128.0K
0.50
/
DeepInfra logoDeepInfra
$0.350$0.400128.0K/128.0K
0.50
/
Hyperbolic logoHyperbolic
$0.400$0.400128.0K/128.0K
0.50
/
Groq logoGroq
$0.590$0.780128.0K/128.0K
0.50
/
Cerebras logoCerebras
$0.600$0.600128.0K/128.0K
0.20
/
Together logoTogether
$0.890$0.890128.0K/128.0K
0.50
/
Fireworks logoFireworks
$0.890$0.890128.0K/128.0K
0.50
/
Bedrock logoBedrock
$0.890$0.890128.0K/128.0K
0.50
/
Sambanova logoSambanova
$5.00$10.00128.0K/128.0K
0.50
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

Llama 3.1 70B Instruct model size

Llama 3.1 70B Instruct has 70 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
70B
15Ttokens
214× tokens-to-params ratio
Large (30–80B)
70B
1B7B70B405B

Llama 3.1 70B Instruct context window

Input and output token limits for Llama 3.1 70B Instruct, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Try now

huggle
Llama 3.1 70B Instructin Huggle

Make it with
Llama 3.1 70B Instruct.

Llama 3.1 70B Instruct

Llama 3.1 70B Instruct latency

Llama 3.1 70B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Llama 3.1 70B Instruct examples

Recent arena outputs from Llama 3.1 70B Instruct, picked from the highest-ranked matchups.

Llama 3.1 70B Instruct license

Llama 3.1 70B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 70.0B parameters.

License
Llama 3.1 Community License
Non-commercial
Parameters
70.0B

Llama 3.1 70B Instruct resources

Official sources for Llama 3.1 70B Instruct: provider documentation, paper or system card, official launch post, source repository, model weights.

Llama 3.1 70B Instruct vs other models

The most-compared alternatives to Llama 3.1 70B Instruct are Qwen3 VL 235B A22B Instruct, Nova Lite, Mistral Large 2. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 3.1 70B Instruct

Models ranked just above and below Llama 3.1 70B Instruct by LLM Stats score.

 

Qwen3 VL 235B A22B Instruct

Score pending
 

Nova Lite

Score pending
 

Mistral Large 2

Score pending
 

Qwen2.5 32B Instruct

Score pending
 

Nova Micro

Score pending
 

Qwen3-Next-80B-A3B-Instruct

Score pending

FAQ

Common questions about Llama 3.1 70B Instruct.

When was Llama 3.1 70B Instruct released?

Llama 3.1 70B Instruct was released on July 23, 2024 by Meta. This is the official Llama 3.1 70B Instruct release date tracked on LLM Stats.

How much does Llama 3.1 70B Instruct cost?

Llama 3.1 70B Instruct pricing starts at $0.20 per million input tokens and $0.20 per million output tokens via Lambda, the lowest price among tracked providers.

How big is Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct has 70 billion parameters. It was trained on 15.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct was created by Meta.

What is the license for Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct is released under the Llama 3.1 Community License license. This is an open-source / open-weight license that permits self-hosting.

What is Llama 3.1 70B Instruct latency?

Llama 3.1 70B Instruct p95 time to first token is 0.20 seconds via Cerebras over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct is available through 9 providers including Lambda, DeepInfra, Hyperbolic, and 6 more.

Where is the Llama 3.1 70B Instruct paper or technical report?

Llama 3.1 70B Instruct has a paper or technical report available at https://ai.meta.com/research/publications/llama-3-open-foundation-and-fine-tuned-chat-models/. Use that source for architecture, training, release and evaluation details.

What models should I compare Llama 3.1 70B Instruct against?

Common Llama 3.1 70B Instruct comparisons include Llama 3.1 70B Instruct vs Qwen3 VL 235B A22B Instruct, Llama 3.1 70B Instruct vs Nova Lite, Llama 3.1 70B Instruct vs Mistral Large 2. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.