Official2026-06-28
DeepSeek V4 API Limits: Context, Pricing and Concurrency
Checked on June 28, 2026: DeepSeek's official Models & Pricing and Rate Limit pages currently show a 1M context window, 384K max output, per-1M token pricing, and account-level concurrency limits of 2500 for DeepSeek V4 Flash and 500 for DeepSeek V4 Pro, with 429s when that cap is exceeded.
What we verified on June 28, 2026
- DeepSeek's official Models & Pricing page currently shows a shared V4 surface with a 1M context length and 384K maximum output for both
deepseek-v4-flashanddeepseek-v4-pro. - The same official page bills in units of per 1M tokens and currently lists
deepseek-v4-flashat$0.0028input cache hit,$0.14input cache miss, and$0.28output, whiledeepseek-v4-prois listed at$0.003625,$0.435, and$0.87respectively. - DeepSeek's official rate-limit page currently treats concurrency as an account-level cap, not a per-key loophole:
deepseek-v4-flashis listed at2500concurrent requests anddeepseek-v4-proat500. - The official rate-limit contract is explicit about failure mode: once the concurrency limit is exceeded, requests receive HTTP
429rather than silently queueing forever. - DeepSeek also documents a capacity-expansion path with no additional cost, but only through a business-needs request workflow. That is a request path, not a guarantee of automatic higher limits.
- The pricing page still warns that product prices may vary, so this item should be read as a current docs snapshot rather than a promise that rates never change.