The AI arena is free today

Open Superagent
DeepSeekReleased on Sep 10, 2026

DeepSeek-V4.1-Flash: Benchmarks, Pricing & Context Window

DeepSeek-V4.1-Flash is a language model from DeepSeek, released in September 2026, with multimodal input, a 1.0M-token context window, and pricing from $0.220/M input, $0.007/M cached input, $0.660/M output.

DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the

Input
TextImage
Output
Text

DeepSeek-V4.1-Flash benchmarks

Capability tiers

Standing within each category, adjusted for leaderboard depth.

Real tasks performance

High-confidence performance for DeepSeek-V4.1-Flash across real-world prompt categories. Only 95% intervals at most 4 points wide are shown.

Performance by conversation depth

How DeepSeek-V4.1-Flash holds up as conversations get longer.

Quality Tracker

DeepSeek-V4.1-Flash Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Thu Sep 10 2026
Notice missing or incorrect data?

DeepSeek-V4.1-Flash pricing

Providers

DeepSeek-V4.1-Flash starts at $0.220 per million input tokens and $0.660 per million output tokens via Fireworks. Reused prompt prefixes cost $0.0070 per million cached input tokens. See all 4 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MCached input $/MOutput $/MContext in / outTTFT p95 sOutput p5 c/sModalities in / out
Fireworks logoFireworks
$0.220$0.0070$0.6601.0M/393.2K
/
DeepInfra logoDeepInfra
$0.300$0.0060$1.201.0M/131.1K
/
DeepSeek logoDeepSeek
$0.300$0.0060$1.201.0M/393.2K
/
Novita logoNovita
$0.300$0.0060$1.201.0M/393.2K
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests.

Loading chart...
Loading chart...
Loading chart...
Loading chart...

DeepSeek-V4.1-Flash model size

DeepSeek-V4.1-Flash has 763.2 billion parameters and was trained on 45 trillion tokens. See how it compares to other models in the same parameter range.

ParametersTraining tokens
763.2BMoE
45Ttokens
59× tokens-to-params ratio
Frontier (200B+)
763.2B
1B7B70B405B

DeepSeek-V4.1-Flash context window

Input and output token limits for DeepSeek-V4.1-Flash, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
393Ktokens
1.6k pages of text
1.0M
8K128K1M

Try now

huggle
DeepSeek-V4.1-Flashin Huggle

Make it with
DeepSeek-V4.1-Flash.

DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash latency

DeepSeek-V4.1-Flash time to first token, sustained output throughput, and failed-request rate from live model usage over the trailing 7 days.

DeepSeek-V4.1-Flash examples

Recent arena outputs from DeepSeek-V4.1-Flash, picked from the highest-ranked matchups.

DeepSeek-V4.1-Flash license

DeepSeek-V4.1-Flash is released under the MIT license, which permits commercial use, has 763.2B parameters.

License
MIT
Commercial use allowed
Parameters
763.2B

MIT License - allows commercial use

DeepSeek-V4.1-Flash resources

Official sources for DeepSeek-V4.1-Flash: provider documentation, official playground, paper or system card, official launch post, source repository, model weights.

DeepSeek-V4.1-Flash vs other models

The most-compared alternatives to DeepSeek-V4.1-Flash are Claude Opus 4.6, GPT-5.2 Pro, ERNIE 5.0. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like DeepSeek-V4.1-Flash

Models ranked just above and below DeepSeek-V4.1-Flash by LLM Stats score.

 

Claude Opus 4.6

Score pending
 

GPT-5.2 Pro

Score pending
 

ERNIE 5.0

Score pending
 

GPT-6 Astra

Score pending
 

Muse Spark 1.3

Score pending
 

Kimi K2.6

Score pending

FAQ

Common questions about DeepSeek-V4.1-Flash.

When was DeepSeek-V4.1-Flash released?

DeepSeek-V4.1-Flash was released on September 10, 2026 by DeepSeek. This is the official DeepSeek-V4.1-Flash release date tracked on LLM Stats.

How much does DeepSeek-V4.1-Flash cost?

DeepSeek-V4.1-Flash pricing starts at $0.22 per million input tokens, $0.01 per million cached input tokens, $0.66 per million output tokens via Fireworks, the lowest price among tracked providers.

How big is DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash has 763.2 billion parameters. It was trained on 45.0 trillion tokens. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash was created by DeepSeek.

What is the license for DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

Is DeepSeek-V4.1-Flash multimodal?

Yes, DeepSeek-V4.1-Flash is multimodal and can accept both text and images as input.

Where can I use DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash is available through 4 providers including Fireworks, DeepInfra, DeepSeek, and 1 more.

Where is the DeepSeek-V4.1-Flash paper or technical report?

DeepSeek-V4.1-Flash has a paper or technical report available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf. Use that source for architecture, training, release and evaluation details.

What models should I compare DeepSeek-V4.1-Flash against?

Common DeepSeek-V4.1-Flash comparisons include DeepSeek-V4.1-Flash vs Claude Opus 4.6, DeepSeek-V4.1-Flash vs GPT-5.2 Pro, DeepSeek-V4.1-Flash vs ERNIE 5.0. Compare them side by side for benchmark scores, pricing, context window, latency and provider availability.