Updated 2026-08-13
DeepSeek V4 API Pricing: Pro, Flash and Major Rivals
DeepSeek V4 remains a strong cost baseline, but its official API price is now time-dependent. From August 17, Flash cache-miss input/output costs $0.22/$0.66 per 1M tokens off-peak and $0.44/$1.32 at peak; Pro costs $0.66/$1.98 off-peak and $1.32/$3.96 at peak. Peak means 09:00–12:00 and 14:00–18:00 Beijing time (01:00–04:00 and 06:00–10:00 UTC); all other hours are off-peak.
Practical verdict
Use DeepSeek V4 as the cost baseline and compare Pro or Flash against the actual workload rather than a vague DeepSeek label. Add paid fallbacks only when measured quality improvements justify the higher cost for a specific request class.
Model snapshot
| Model | Provider | Strengths | Context | Cost signal |
|---|---|---|---|---|
| DeepSeek V4 | DeepSeek | Coding, Long Context, Cost-Efficiency | 1M | $0.32 / 1M avg tokens |
| GPT 5.4 | OpenAI | Reasoning, Tool Calling, Multimodal | 1M | $8.75 / 1M avg tokens |
| Claude Sonnet 4.7 | Anthropic | Coding, Agentic, Long Context | 1M | $9.00 / 1M avg tokens |
| Gemini 3.1 Pro | Reasoning, Multimodal, Long Context | 2M | $7.00 / 1M avg tokens | |
| Qwen 3.5 | Alibaba | Multilingual, Reasoning, Open Source, Cost-Efficiency | 1M | $1.14 / 1M avg tokens |
| MiniMax M2.7 | MiniMax | Agentic, Coding, Long Context, Cost-Efficiency | 205K | $0.75 / 1M avg tokens |
| GLM 5 | Zhipu AI | Coding, Agentic, Multilingual, Cost-Efficiency | 200K | $0.90 / 1M avg tokens |
Cost signals are comparison data used by this site. Verify live provider pricing before production purchasing decisions.
Use-case routing table
| Use case | DeepSeek fit | Alternative fit | Decision note |
|---|---|---|---|
| Default chat and coding API | Best cost baseline | Premium fallback | Start with DeepSeek before paying premium-provider prices on every request. |
| Long-context research | Strong on 1M context | Gemini/Claude strong | Large multimodal inputs can still justify specialized models. |
| Multilingual production | Strong | Qwen/GLM strong | Cost and native-language quality both matter in real deployments. |
| Interactive experience product | Good | MiniMax/Grok strong | Experience quality can justify a different default. |
How to read pricing comparisons in the V4 era
Token price is only the starting point. Compare total cost per successful task, including input tokens, output tokens, failed calls, retries, human correction, and the time window. From August 17, DeepSeek V4 Pro cache-miss input/output costs $0.66/$1.98 per 1M tokens off-peak and $1.32/$3.96 at peak. The right comparison is often V4-Flash versus premium defaults, or V4-Pro versus premium review routes, not simply DeepSeek versus everyone else.
Why DeepSeek V4 is the baseline
A DeepSeek-first pricing page gives buyers a concrete anchor: what quality can they get from an official 1M-context flagship before paying premium model prices? That anchor is what turns comparison traffic into pricing-page intent.
Why Flash deserves separate attention
DeepSeek V4 Flash should not be treated as a weak budget footnote. Its quality is excellent for everyday coding, chat, retrieval, and repeated tool steps, while keeping the same 1M-context story. From August 17, Flash cache-miss input/output costs $0.22/$0.66 per 1M tokens off-peak and $0.44/$1.32 at peak, with cache-hit input at $0.007 off-peak or $0.014 at peak.
Where the pricing page fits
The pricing page should only show plans backed by inventory. This comparison page can mention many models, but it should route purchase intent to the actual in-stock DeepSeek-led Coding Plans.
FAQ
Is DeepSeek V4 the cheapest AI API?
DeepSeek V4 is one of the most cost-efficient options for many developer workflows. From August 17, Flash starts at $0.007 cache-hit input, $0.22 cache-miss input, and $0.66 output per 1M tokens off-peak; Pro starts at $0.022, $0.66, and $1.98. Peak rates are higher, so compare the schedule and verify live pricing before estimating production spend.
Is DeepSeek V4 Flash good enough for production?
Yes for many high-volume routes. Flash is excellent for routine coding, chat, retrieval, and tool calls; keep Pro for harder reasoning and review-heavy turns.
What should I compare besides token price?
Compare latency, retries, correctness, context fit, and cost per accepted output.
Why are some compared models not on the pricing page?
Comparison coverage is independent from inventory. Only in-stock Coding Plan products are purchasable.