The AI arena is free today

Open Superagent
DeepSeekReleased on Dec 25, 2024

DeepSeek-V3: Benchmarks, Pricing & Context Window

DeepSeek-V3 is a language model from DeepSeek, released in December 2024, with a 164K-token context window, and pricing from $0.320/M input and $0.890/M output.

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on

Input
Text
Output
Text

DeepSeek-V3 benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for DeepSeek-V3 across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How DeepSeek-V3 holds up as conversations get longer.

Quality Tracker

DeepSeek-V3 Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Oct 08 2026
Notice missing or incorrect data?

DeepSeek-V3 pricing

Providers

DeepSeek-V3 starts at $0.270 per million input tokens and $1.10 per million output tokens via DeepSeek. See all 2 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
DeepSeek logoDeepSeek
$0.270—$1.10131.1K/131.1K
0.50
—
/
DeepInfra logoDeepInfra
$0.320—$0.890163.8K/163.8K
—
—
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...

DeepSeek-V3 model size

DeepSeek-V3 has 671 billion parameters and was trained on 14.8 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
671B
14.8Ttokens
22× tokens-to-params ratio
Frontier (200B+)
671B
1B7B70B405B

DeepSeek-V3 context window

Input and output token limits for DeepSeek-V3, plus how it ranks on long-context understanding.

InputOutput
164Ktokens
164Ktokens
≈ 246 pages of text
164K
8K128K1M

Try now

huggle
DeepSeek-V3in Huggle

Make it with
DeepSeek-V3.

DeepSeek-V3

DeepSeek-V3 latency

DeepSeek-V3 time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

DeepSeek-V3 examples

Recent arena outputs from DeepSeek-V3, picked from the highest-ranked matchups.

DeepSeek-V3 license

DeepSeek-V3 is released under the MIT + Model License (Commercial use allowed) license, which restricts commercial use, has 671.0B parameters.

License
MIT + Model License (Commercial use allowed)
Non-commercial
Parameters
671.0B

DeepSeek-V3 resources

Official sources for DeepSeek-V3: provider documentation, official playground, paper or system card, source repository, model weights.

DeepSeek-V3 vs other models

The most-compared alternatives to DeepSeek-V3 are Claude 3.5 Sonnet, Phi 4 Reasoning Plus, GPT-4o. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like DeepSeek-V3

Models ranked just above and below DeepSeek-V3 by LLM Stats score.

 

Claude 3.5 Sonnet

Score pending
 

Phi 4 Reasoning Plus

Score pending
 

GPT-4o

Score pending
 

Grok-2

Score pending
 

Gemini 1.5 Pro

Score pending
 

Qwen3 VL 32B Thinking

Score pending

FAQ

Common questions about DeepSeek-V3.

When was DeepSeek-V3 released?

DeepSeek-V3 was released on December 25, 2024 by DeepSeek. This is the official DeepSeek-V3 release date tracked on LLM Stats.

How much does DeepSeek-V3 cost?

DeepSeek-V3 pricing starts at $0.27 per million input tokens and $1.10 per million output tokens via DeepSeek, the lowest price among tracked providers.

How big is DeepSeek-V3?

DeepSeek-V3 has 671 billion parameters. It was trained on 14.8 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created DeepSeek-V3?

DeepSeek-V3 was created by DeepSeek.

What is the license for DeepSeek-V3?

DeepSeek-V3 is released under the MIT + Model License (Commercial use allowed) license. This is an open-source / open-weight license that permits self-hosting.

What is DeepSeek-V3 latency?

DeepSeek-V3 p95 time to first token is 0.50 seconds via DeepSeek over the trailing 7 days. Lower time to first token means the model begins responding sooner for chat, agents and model workloads.

Where can I use DeepSeek-V3?

DeepSeek-V3 is available through 2 providers including DeepSeek, DeepInfra.

Where is the DeepSeek-V3 paper or technical report?

DeepSeek-V3 has a paper or technical report available at https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSeek_V3.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare DeepSeek-V3 against?

Common DeepSeek-V3 comparisons include DeepSeek-V3 vs Claude 3.5 Sonnet, DeepSeek-V3 vs Phi 4 Reasoning Plus, DeepSeek-V3 vs GPT-4o. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.