DeepSeekReleased on Jul 31, 2026

DeepSeek-V4-Flash-0731: API Pricing, Context Window & Benchmarks

DeepSeek-V4-Flash-0731 is a language model from DeepSeek, released in July 2026, with a 1.0M-token context window, and pricing from $0.090/M input and $0.180/M output.

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version with substantially stronger agentic capabilities. It retains the V4-Flash architecture and adds an attached DSpark speculative-decoding

Input
Text
Output
Text

DeepSeek-V4-Flash-0731 benchmarks

Rankings

Quality Tracker

DeepSeek-V4-Flash-0731 Performance Across Datasets

Scores sourced from the model's scorecard, paper, or official blog posts

LLM Stats Logollm-stats.com - Mon Aug 03 2026
Notice missing or incorrect data?

DeepSeek-V4-Flash-0731 pricing

Providers

DeepSeek-V4-Flash-0731 starts at $0.0900 per million input tokens and $0.180 per million output tokens via DeepInfra. See all 3 providers below with their per-token pricing, latency, throughput, and modality support.

ProviderInput $/MOutput $/MContext in / outTTFT p50 / p95 sOutput avg / p5 c/sSuccess 7dModalities in / out
DeepInfra logoDeepInfra
$0.0900$0.1801.0M/65.5K
/
/
/
Fireworks logoFireworks
$0.140$0.2801.0M/65.5K
/
/
/
Novita logoNovita
$0.140$0.2801.0M/393.2K
/
/
/

Cached input is the discounted price for prompt tokens served from a provider cache. TTFT is time to first token. Output is characters per second; p5 is the sustained floor exceeded by 95% of observed requests. Success is calculated from completed versus failed requests over the trailing seven days.

Loading chart...
Loading chart...
Loading chart...

DeepSeek-V4-Flash-0731 model size

DeepSeek-V4-Flash-0731 has 304 billion parameters. See how it compares to other models in the same parameter range.

Parameters
304BMoE
Frontier (200B+)
304B
1B7B70B405B

DeepSeek-V4-Flash-0731 context window

Input and output token limits for DeepSeek-V4-Flash-0731, plus how it ranks on long-context understanding.

InputOutput
1.0Mtokens
393Ktokens
1.6k pages of text
1.0M
8K128K1M

DeepSeek-V4-Flash-0731 API

POST/v1/chat/completions

Run a request to see the response

Use it in your code

Billed at $0.09 input / $0.18 output per 1M tokens through the LLM Stats gateway.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://gateway.llm-stats.com/v1"
)

response = client.chat.completions.create(
    model="deepseek-v4-flash-0731",
    messages=[
        {"role": "user", "content": "What is machine learning?"}
    ]
)

print(response.choices[0].message.content)

Need an API key? Create one above in the playground, or read the API documentation.

DeepSeek-V4-Flash-0731 latency

DeepSeek-V4-Flash-0731 time to first token, sustained output throughput, and failed-request rate from live API traffic over the trailing 7 days.

DeepSeek-V4-Flash-0731 examples

Recent arena outputs from DeepSeek-V4-Flash-0731, picked from the highest-ranked matchups.

DeepSeek-V4-Flash-0731 license

DeepSeek-V4-Flash-0731 is released under the MIT license, which permits commercial use, has 304.0B parameters.

License
MIT
Commercial use allowed
Parameters
304.0B

MIT License - allows commercial use

DeepSeek-V4-Flash-0731 resources

Official sources for DeepSeek-V4-Flash-0731: api documentation, official playground, paper or system card, source repository, model weights.

DeepSeek-V4-Flash-0731 vs other models

The most-compared alternatives to DeepSeek-V4-Flash-0731 are Claude Opus 4.6, Qwen3.7 Max, Claude Opus 4.8. Open any pair side-by-side for benchmarks, pricing, context, and latency.

Models like DeepSeek-V4-Flash-0731

Models ranked just above and below DeepSeek-V4-Flash-0731 by LLM Stats score.

 

Claude Opus 4.6

Score pending
 

Qwen3.7 Max

Score pending
 

Claude Opus 4.8

Score pending
 

Muse Spark 1.1

Score pending
 

Claude Opus 4.7

Score pending
 

Kimi K3

Score pending

FAQ

Common questions about DeepSeek-V4-Flash-0731.

When was DeepSeek-V4-Flash-0731 released?

DeepSeek-V4-Flash-0731 was released on July 31, 2026 by DeepSeek. This is the official DeepSeek-V4-Flash-0731 release date tracked on LLM Stats.

How much does DeepSeek-V4-Flash-0731 cost?

DeepSeek-V4-Flash-0731 costs $0.09 per million input tokens and $0.18 per million output tokens through the LLM Stats API, which works with any OpenAI-compatible SDK. Across tracked providers, the lowest price is $0.09 per million input tokens via DeepInfra.

Is DeepSeek-V4-Flash-0731 available via API?

Yes. DeepSeek-V4-Flash-0731 is available through the LLM Stats API and works with any OpenAI-compatible SDK — point your client at the gateway base URL and pass the model name. It is served by 3 providers tracked on LLM Stats.

How big is DeepSeek-V4-Flash-0731?

DeepSeek-V4-Flash-0731 has 304 billion parameters. It ships as an open-weight model, so you can download and run it on your own hardware.

Who created DeepSeek-V4-Flash-0731?

DeepSeek-V4-Flash-0731 was created by DeepSeek.

What is the license for DeepSeek-V4-Flash-0731?

DeepSeek-V4-Flash-0731 is released under the MIT license. This is an open-source / open-weight license that permits self-hosting.

Where can I use DeepSeek-V4-Flash-0731?

DeepSeek-V4-Flash-0731 is available through 3 providers including DeepInfra, Fireworks, Novita.

Where is the DeepSeek-V4-Flash-0731 paper or technical report?

DeepSeek-V4-Flash-0731 has a paper or technical report available at https://arxiv.org/abs/2606.19348. Use that source for architecture, training, release and evaluation details.

What models should I compare DeepSeek-V4-Flash-0731 against?

Common DeepSeek-V4-Flash-0731 comparisons include DeepSeek-V4-Flash-0731 vs Claude Opus 4.6, DeepSeek-V4-Flash-0731 vs Qwen3.7 Max, DeepSeek-V4-Flash-0731 vs Claude Opus 4.8. Compare them side by side for benchmark scores, pricing, context window, latency and API availability.