The AI arena is free today

Open Superagent
MetaReleased on Jul 23, 2024

Llama 3.1 8B Instruct: Benchmarks, Pricing & Context Window

Llama 3.1 8B Instruct is a language model from Meta, released in July 2024, with a 131K-token context window, and pricing from $0.020/M input and $0.040/M output.

Llama 3.1 8B Instruct is a multilingual large language model optimized for dialogue use cases. It features a 128K context length, state-of-the-art tool use, and strong reasoning capabilities.

Input
Text
Output
Text

Llama 3.1 8B Instruct benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for Llama 3.1 8B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How Llama 3.1 8B Instruct holds up as conversations get longer.

Quality Tracker

Llama 3.1 8B Instruct Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Oct 08 2026
Notice missing or incorrect data?

Llama 3.1 8B Instruct pricing

Providers

Llama 3.1 8B Instruct starts at $0.0200 per million input tokens and $0.0400 per million output tokens via DeepInfra. See all 9 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.0200—$0.0400131.1K/131.1K
—
—
/
Lambda logoLambda
$0.0300—$0.0300131.1K/131.1K
0.50
—
/
Groq logoGroq
$0.0500—$0.0800131.1K/131.1K
0.50
—
/
Sambanova logoSambanova
$0.100—$0.200131.1K/131.1K
0.50
—
/
Cerebras logoCerebras
$0.100—$0.100131.1K/131.1K
0.20
—
/
Hyperbolic logoHyperbolic
$0.100—$0.100131.1K/131.1K
0.50
—
/
Together logoTogether
$0.200—$0.200131.1K/131.1K
0.50
—
/
Fireworks logoFireworks
$0.200—$0.200131.1K/131.1K
0.50
—
/
Bedrock logoBedrock
$0.220—$0.220131.1K/131.1K
0.50
—
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

Llama 3.1 8B Instruct model size

Llama 3.1 8B Instruct has 8 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
8B
15Ttokens
1875× tokens-to-params ratio
Small (3–10B)
8B
1B7B70B405B

Llama 3.1 8B Instruct context window

Input and output token limits for Llama 3.1 8B Instruct, plus how it ranks on long-context understanding.

InputOutput
131Ktokens
131Ktokens
≈ 197 pages of text
131K
8K128K1M

Try now

huggle
Llama 3.1 8B Instructin Huggle

Make it with
Llama 3.1 8B Instruct.

Llama 3.1 8B Instruct

Llama 3.1 8B Instruct latency

Llama 3.1 8B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

Llama 3.1 8B Instruct examples

Recent arena outputs from Llama 3.1 8B Instruct, picked from the highest-ranked matchups.

Llama 3.1 8B Instruct license

Llama 3.1 8B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 8.0B parameters, has a knowledge cutoff of December 2023.

License
Llama 3.1 Community License
Non-commercial
Parameters
8.0B
Knowledge cutoff
December 2023

Llama 3.1 8B Instruct resources

Official sources for Llama 3.1 8B Instruct: provider documentation, official launch post, source repository, model weights.

Llama 3.1 8B Instruct vs other models

The most-compared alternatives to Llama 3.1 8B Instruct are Qwen3 VL 30B A3B Thinking, Hermes 3 70B, Qwen2.5-Coder 32B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like Llama 3.1 8B Instruct

Models ranked just above and below Llama 3.1 8B Instruct by LLM Stats score.

 

Qwen3 VL 30B A3B Thinking

Score pending
 

Hermes 3 70B

Score pending
 

Qwen2.5-Coder 32B Instruct

Score pending
 

Gemini 1.0 Pro

Score pending
 

Phi-3.5-mini-instruct

Score pending
 

Jamba 1.5 Mini

Score pending

FAQ

Common questions about Llama 3.1 8B Instruct.

When was Llama 3.1 8B Instruct released?

Llama 3.1 8B Instruct was released on July 23, 2024 by Meta. This is the official Llama 3.1 8B Instruct release date tracked on LLM Stats.

How much does Llama 3.1 8B Instruct cost?

Llama 3.1 8B Instruct pricing starts at $0.02 per million input tokens and $0.04 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct has 8 billion parameters. It was trained on 15.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct was created by Meta.

What is the license for Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is released under the Llama 3.1 Community License license. This is an open-source / open-weight license that permits self-hosting.

What is the knowledge cutoff date for Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct has a knowledge cutoff of December 2023, meaning it was trained on data up to that point and may not know about events after it.

What is Llama 3.1 8B Instruct latency?

Llama 3.1 8B Instruct p95 time to first token is 0.20 seconds via Cerebras over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is available through 9 providers including DeepInfra, Lambda, Groq, and 6 more.

Where is the Llama 3.1 8B Instruct paper or technical report?

Llama 3.1 8B Instruct has a paper or technical report available at https://ai.meta.com/blog/meta-llama-3-1/. Use that source for architecture, training, release and evaluation details.

What models should I compare Llama 3.1 8B Instruct against?

Common Llama 3.1 8B Instruct comparisons include Llama 3.1 8B Instruct vs Qwen3 VL 30B A3B Thinking, Llama 3.1 8B Instruct vs Hermes 3 70B, Llama 3.1 8B Instruct vs Qwen2.5-Coder 32B Instruct. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.