- Organizations
- Meta
- Llama 3.1 8B Instruct
Llama 3.1 8B Instruct: Benchmarks, Pricing & Context Window
Llama 3.1 8B Instruct is a language model from Meta, released in July 2024, with a 131K-token context window, and pricing from $0.020/M input and $0.040/M output.
Llama 3.1 8B Instruct is a multilingual large language model optimized for dialogue use cases. It features a 128K context length, state-of-the-art tool use, and strong reasoning capabilities.
Llama 3.1 8B Instruct benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for Llama 3.1 8B Instruct across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How Llama 3.1 8B Instruct holds up as conversations get longer.
Quality Tracker
Llama 3.1 8B Instruct Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
Llama 3.1 8B Instruct pricing
Providers
Llama 3.1 8B Instruct starts at $0.0200 per million input tokens and $0.0400 per million output tokens via DeepInfra. See all 9 providers below with their per-token pricing, latency, throughput, and modality support.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.0200 | — | $0.0400 | 131.1K/131.1K | — | — | / | |
| $0.0300 | — | $0.0300 | 131.1K/131.1K | 0.50 | — | / | |
| $0.0500 | — | $0.0800 | 131.1K/131.1K | 0.50 | — | / | |
| $0.100 | — | $0.200 | 131.1K/131.1K | 0.50 | — | / | |
| $0.100 | — | $0.100 | 131.1K/131.1K | 0.20 | — | / | |
| $0.100 | — | $0.100 | 131.1K/131.1K | 0.50 | — | / | |
| $0.200 | — | $0.200 | 131.1K/131.1K | 0.50 | — | / | |
| $0.200 | — | $0.200 | 131.1K/131.1K | 0.50 | — | / | |
| $0.220 | — | $0.220 | 131.1K/131.1K | 0.50 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
Llama 3.1 8B Instruct model size
Llama 3.1 8B Instruct has 8 billion parameters and was trained on 15 trillion tokens. See how it compares to other models in the same parameter range.
Llama 3.1 8B Instruct context window
Input and output token limits for Llama 3.1 8B Instruct, plus how it ranks on long-context understanding.
Try now
Make it with
Llama 3.1 8B Instruct.
Llama 3.1 8B Instruct latency
Llama 3.1 8B Instruct time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
Llama 3.1 8B Instruct examples
Recent arena outputs from Llama 3.1 8B Instruct, picked from the highest-ranked matchups.
Llama 3.1 8B Instruct license
Llama 3.1 8B Instruct is released under the Llama 3.1 Community License license, which restricts commercial use, has 8.0B parameters, has a knowledge cutoff of December 2023.
- License
- Llama 3.1 Community License
- Non-commercial
- Parameters
- 8.0B
- Knowledge cutoff
- December 2023
Llama 3.1 8B Instruct resources
Official sources for Llama 3.1 8B Instruct: provider documentation, official launch post, source repository, model weights.
Llama 3.1 8B Instruct vs other models
The most-compared alternatives to Llama 3.1 8B Instruct are Qwen3 VL 30B A3B Thinking, Hermes 3 70B, Qwen2.5-Coder 32B Instruct. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like Llama 3.1 8B Instruct
Models ranked just above and below Llama 3.1 8B Instruct by LLM Stats score.
FAQ
Common questions about Llama 3.1 8B Instruct.