- Organizations
- DeepSeek
- DeepSeek R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B: Benchmarks, Pricing & Context Window
DeepSeek R1 Distill Llama 70B is a language model from DeepSeek, released in January 2025.
DeepSeek-R1 is the first-generation reasoning model built atop DeepSeek-V3 (671B total parameters, 37B activated per token). It incorporates large-scale reinforcement learning (RL) to enhance its chain-of-thought and reasoning
DeepSeek R1 Distill Llama 70B benchmarks
Capability tiers
Standing within each category, adjusted for leaderboard depth.
Real tasks performance
High-confidence performance for DeepSeek R1 Distill Llama 70B across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.
Performance by conversation depth
How DeepSeek R1 Distill Llama 70B holds up as conversations get longer.
Quality Tracker
DeepSeek R1 Distill Llama 70B Performance Across Datasets
Scores sourced from the model's scorecard, paper, or official blog posts
DeepSeek R1 Distill Llama 70B pricing
Providers
DeepSeek R1 Distill Llama 70B starts at $0.100 per million input tokens and $0.400 per million output tokens via DeepInfra.
| Provider | Input $/M | Cached input $/M | Output $/M | Context in / out | TTFT p95 s | Output p5 c/s | Modalities in / out |
|---|---|---|---|---|---|---|---|
| $0.100 | — | $0.400 | 128.0K/128.0K | 0.65 | — | / |
Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.
DeepSeek R1 Distill Llama 70B model size
DeepSeek R1 Distill Llama 70B has 70.6 billion parameters and was trained on 14.8 trillion tokens. See how it compares to other models in the same parameter range.
DeepSeek R1 Distill Llama 70B context window
Input and output token limits for DeepSeek R1 Distill Llama 70B, plus how it ranks on long-context understanding.
Try now
Make it with
DeepSeek R1 Distill Llama 70B.
DeepSeek R1 Distill Llama 70B latency
DeepSeek R1 Distill Llama 70B time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.
DeepSeek R1 Distill Llama 70B examples
Recent arena outputs from DeepSeek R1 Distill Llama 70B, picked from the highest-ranked matchups.
DeepSeek R1 Distill Llama 70B license
DeepSeek R1 Distill Llama 70B is released under the MIT license, which permits commercial use, has 70.6B parameters.
- License
- MIT
- Commercial use allowed
- Parameters
- 70.6B
MIT License - allows commercial use
DeepSeek R1 Distill Llama 70B resources
Official sources for DeepSeek R1 Distill Llama 70B: provider documentation, official playground, paper or system card, source repository, model weights.
DeepSeek R1 Distill Llama 70B vs other models
The most-compared alternatives to DeepSeek R1 Distill Llama 70B are o1-pro, DeepSeek R1 Zero, Phi 4 Reasoning. Open any pair side-by-side for benchmarks, pricing, context, and latency.
Models like DeepSeek R1 Distill Llama 70B
Models ranked just above and below DeepSeek R1 Distill Llama 70B by LLM Stats score.
FAQ
Common questions about DeepSeek R1 Distill Llama 70B.