DeepSeekReleased on Jan 20, 2025

DeepSeek R1 Distill Llama 70B: API Pricing, Context Window & Benchmarks

DeepSeek R1 Distill Llama 70B is a language model from DeepSeek, released in January 2025.

DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning

Input
Text
Output
Text

DeepSeek R1 Distill Llama 70B benchmarks

Rankings

Quality Tracker

DeepSeek R1 Distill Llama 70B Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Tue Jul 21 2026
Notice missing or incorrect data?

DeepSeek R1 Distill Llama 70B pricing

Providers

DeepSeek R1 Distill Llama 70B starts at $0.100 per million input tokens and $0.400 per million output tokens via DeepInfra.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.100$0.400128.0K/128.0K
/0.65
37/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

DeepSeek R1 Distill Llama 70B model size

DeepSeek R1 Distill Llama 70B has 70.6 billion parameters and was trained on 14.8 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
70.6B
14.8Ttokens
210× tokens-to-params ratio
Large (30–80B)
70.6B
1B7B70B405B

DeepSeek R1 Distill Llama 70B context window

Input and output token limits for DeepSeek R1 Distill Llama 70B, plus how it ranks on long-context understanding.

InputOutput
128Ktokens
128Ktokens
192 pages of text
128K
8K128K1M

DeepSeek R1 Distill Llama 70B API

Available from the model provider

DeepSeek R1 Distill Llama 70B has an official provider API. It is not currently routed through the LLM Stats gateway.

Read the official API documentation

DeepSeek R1 Distill Llama 70B latency

DeepSeek R1 Distill Llama 70B time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

DeepSeek R1 Distill Llama 70B examples

Recent arena outputs from DeepSeek R1 Distill Llama 70B, picked from the highest-ranked matchups.

DeepSeek R1 Distill Llama 70B license

DeepSeek R1 Distill Llama 70B is released under the MIT license, which permits commercial use, has 70.6B parameters.

License
MIT
Commercial use allowed
Parameters
70.6B

MIT License - allows commercial use

DeepSeek R1 Distill Llama 70B resources

Official sources for DeepSeek R1 Distill Llama 70B: api documentation, official playground, paper or system card, source repository, model weights.

DeepSeek R1 Distill Llama 70B vs other models

The most-compared alternatives to DeepSeek R1 Distill Llama 70B are DeepSeek R1 Zero, Phi 4 Reasoning, QwQ-32B. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like DeepSeek R1 Distill Llama 70B

Models ranked just above and below DeepSeek R1 Distill Llama 70B by LLM Stats score.

 

DeepSeek R1 Zero

Score pending
 

Phi 4 Reasoning

Score pending
 

QwQ-32B

Score pending
 

DeepSeek R1 Distill Qwen 32B

Score pending
 

Qwen3 30B A3B

Score pending
 

Ministral 3 (8B Reasoning 2512)

Score pending

FAQ

Common questions about DeepSeek R1 Distill Llama 70B.

When was DeepSeek R1 Distill Llama 70B released?

DeepSeek R1 Distill Llama 70B was released on January 20, 2025 by DeepSeek. This is the official DeepSeek R1 Distill Llama 70B release date tracked on LLM Stats.

How much does DeepSeek R1 Distill Llama 70B cost?

DeepSeek R1 Distill Llama 70B pricing starts at $0.10 per million input tokens and $0.40 per million output tokens via DeepInfra, the lowest price among tracked providers.

Is DeepSeek R1 Distill Llama 70B available via API?

Yes, DeepSeek R1 Distill Llama 70B is available via API. See the official documentation for authentication and endpoint details. It is served by 1 provider tracked on LLM Stats.

How big is DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B has 70.6 billion parameters. It was trained on 14.8 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B was created by DeepSeek.

What is the license for DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

What is DeepSeek R1 Distill Llama 70B latency?

DeepSeek R1 Distill Llama 70B p95 time to first token is 0.65 seconds via DeepInfra over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and API workloads.

Where can I use DeepSeek R1 Distill Llama 70B?

DeepSeek R1 Distill Llama 70B is available through 1 provider including DeepInfra.

Where is the DeepSeek R1 Distill Llama 70B paper or technical report?

DeepSeek R1 Distill Llama 70B has a paper or technical report available at https://arxiv.org/pdf/2501.12948. Use that source for architecture, training, release and evaluation details.

What models should I compare DeepSeek R1 Distill Llama 70B against?

Common DeepSeek R1 Distill Llama 70B comparisons include DeepSeek R1 Distill Llama 70B vs DeepSeek R1 Zero, DeepSeek R1 Distill Llama 70B vs Phi 4 Reasoning, DeepSeek R1 Distill Llama 70B vs QwQ-32B. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.