The AI arena is free today

Open Superagent
DeepSeekReleased on Jan 20, 2025

DeepSeek R1 Distill Llama 70B: Benchmarks, Pricing & Context Window

DeepSeek R1 Distill Llama 70B is a language model from DeepSeek, released in January 2025.

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning

Input
Text
Output
Text

DeepSeek R1 Distill Llama 70B benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for DeepSeek R1 Distill Llama 70B across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How DeepSeek R1 Distill Llama 70B holds up as conversations get longer.

Quality Tracker

DeepSeek R1 Distill Llama 70B Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Sat Sep 19 2026
Notice missing or incorrect data?

DeepSeek R1 Distill Llama 70B pricing

Providers

DeepSeek R1 Distill Llama 70B starts at $0.100 per million input tokens and $0.400 per million output tokens via DeepInfra.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepInfra logoDeepInfra
$0.100$0.400128.0K/128.0K
0.65
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

DeepSeek R1 Distill Llama 70B model size

DeepSeek R1 Distill Llama 70B has 70.6 billion parameters and was trained on 14.8 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
70.6B
14.8Ttokens
210× tokens-to-params ratio
Large (30–80B)
70.6B
1B7B70B405B

DeepSeek R1 Distill Llama 70B context window

Input and output token limits for DeepSeek R1 Distill Llama 70B, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

Try now

huggle
DeepSeek R1 Distill Llama 70Bin Huggle

Make it with
DeepSeek R1 Distill Llama 70B.

DeepSeek R1 Distill Llama 70B

DeepSeek R1 Distill Llama 70B latency

DeepSeek R1 Distill Llama 70B time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

DeepSeek R1 Distill Llama 70B examples

Recent arena outputs from DeepSeek R1 Distill Llama 70B, picked from the highest-ranked matchups.

DeepSeek R1 Distill Llama 70B license

DeepSeek R1 Distill Llama 70B is released under the MIT license, which permits commercial use, has 70.6B parameters.

License
MIT
Commercial use allowed
Parameters
70.6B

MIT License - allows commercial use

DeepSeek R1 Distill Llama 70B resources

Official sources for DeepSeek R1 Distill Llama 70B: provider documentation, official playground, paper or system card, source repository, model weights.

DeepSeek R1 Distill Llama 70B vs other models

The most-compared alternatives to DeepSeek R1 Distill Llama 70B are o1-pro, DeepSeek R1 Zero, Phi 4 Reasoning. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like DeepSeek R1 Distill Llama 70B

Models ranked just above and below DeepSeek R1 Distill Llama 70B by LLM Stats score.

 

o1-pro

Score pending
 

DeepSeek R1 Zero

Score pending
 

Phi 4 Reasoning

Score pending
 

QwQ-32B

Score pending
 

DeepSeek R1 Distill Qwen 32B

Score pending
 

Qwen3 30B A3B

Score pending

FAQ

Common questions about DeepSeek R1 Distill Llama 70B.

When was DeepSeek R1 Distill Llama 70B released?

DeepSeek R1 Distill Llama 70B was released on January 20, 2025 by DeepSeek. This is the official DeepSeek R1 Distill Llama 70B release date tracked on LLM Stats.

How much does DeepSeek R1 Distill Llama 70B cost?

DeepSeek R1 Distill Llama 70B pricing starts at $0.10 per million input tokens and $0.40 per million output tokens via DeepInfra, the lowest price among tracked providers.

How big is DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B has 70.6 billion parameters. It was trained on 14.8 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B was created by DeepSeek.

What is the license for DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is DeepSeek R1 Distill Llama 70B latency?

DeepSeek R1 Distill Llama 70B p95 time to first token is 0.65 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B is available through 1 provider including DeepInfra.

Where is the DeepSeek R1 Distill Llama 70B paper or technical report?

DeepSeek R1 Distill Llama 70B has a paper or technical report available at https://arxiv.org/pdf/2501.12948. Use that source for architecture, training, release and evaluation details.

What models should I compare DeepSeek R1 Distill Llama 70B against?

Common DeepSeek R1 Distill Llama 70B comparisons include DeepSeek R1 Distill Llama 70B vs o1-pro, DeepSeek R1 Distill Llama 70B vs DeepSeek R1 Zero, DeepSeek R1 Distill Llama 70B vs Phi 4 Reasoning. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.