DeepSeek

DeepSeek pricing charges per million tokens and splits input by whether the token hit the disk cache.

Pricing Model:

Pricing Model:

Pure usage, prepaid

Pure usage, prepaid

Pure usage, prepaid

Packaging Model:

Packaging Model:

Two tiers

Two tiers

Two tiers

Credit Model:

Credit Model:

Granted balance and topped-up balance, granted spent first

Granted balance and topped-up balance, granted spent first

Granted balance and topped-up balance, granted spent first

Updated on:

DeepSeek pricing: the cache decides what an input token costs

DeepSeek pricing charges per million tokens and splits input by whether the token hit the disk cache. A cache hit on deepseek-flash costs $0.003 per million off-peak, a cache miss costs $0.15, and that's a 50x gap on the same token. Every rate then doubles during peak hours. Nothing else in this index prices the lookup outcome rather than the act of caching, so what you pay depends on something you don't directly control.

Key takeaways

  • DeepSeek charges cache-hit and cache-miss input at separate rates, 50x apart on deepseek-flash and 30x apart on deepseek-v4-pro.

  • Peak hours run 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and double every rate. That's 35 hours of a 168-hour week, so off-peak applies roughly 79% of the time.

  • Caching is on by default with no parameter to set, and DeepSeek charges nothing to write a cache entry.

  • There's no plan, no allowance and no overage. Calls return HTTP 402 the moment the prepaid balance hits zero.

DeepSeek pricing in 2026

Model

Cache hit input

Cache miss input

Output

Concurrency

 

deepseek-flash (V4.1 Flash)

$0.003 / $0.006

$0.15 / $0.30

$0.60 / $1.20

2,500

deepseek-v4-pro (V4-Pro-0813)

$0.022 / $0.044

$0.66 / $1.32

$1.98 / $3.96

500

Source: api-docs.deepseek.com/quick_start/pricing, read 22 September 2026.

What DeepSeek actually meters

DeepSeek meters tokens, and it's the only vendor in this index that prices the outcome of a cache lookup rather than the work of building the cache. OpenAI and Anthropic both charge a premium to write a cache entry and a discount to read one. DeepSeek charges nothing to write and applies one of two input rates depending on whether your prefix matched.

Every call returns prompt_cache_hit_tokens and prompt_cache_miss_tokens in its usage block, so the split is auditable per request rather than inferred from the invoice.

The gap is severe enough to dominate a bill. On deepseek-flash a cache hit costs 2% of a cache miss. On deepseek-v4-pro it costs roughly 3.3%. Output is the milder multiple: 4x cache-miss input on Flash and 3x on Pro, a narrower spread than the 5x most Western providers run.

Caching itself is automatic, enabled for all users by default with no parameter to toggle. Hits require a full match against a persisted prefix unit, and DeepSeek persists those at request boundaries, at common prefixes detected across requests, and at fixed token intervals inside long inputs. The docs are explicit that this is best effort with no guaranteed hit rate, and that an unused cache clears within hours to days.

How DeepSeek's balance works

DeepSeek runs prepaid, with two balances per account. A granted balance covers promotional credit and a topped-up balance covers what you've paid in. Charges draw from the granted balance first when both exist, which is the deduction order you'd want.

Granted credit expires. The balance endpoint returns granted_balance as "the total not expired granted balance", which confirms expiry exists, but the docs never publish the window or what triggers a grant. That's the one real gap in an otherwise complete rate card.

There are no credit packs, rollover rules or pooling to document, because DeepSeek doesn't sell credits as a product. The balance is a cash float against a meter. A user_id parameter isolates content safety, cache and scheduling per end user, but it doesn't split the balance.

What happens when you hit the limit

Nothing throttles and nothing auto-charges. At zero balance the API returns HTTP 402 Insufficient Balance and you top up to resume, which is a hard stop rather than an invoice surprise.

The separate ceiling is concurrency: 2,500 in-flight requests on Flash and 500 on Pro, counted per account across all API keys. Exceeding it returns HTTP 429. DeepSeek grants expansion on request and states plainly that it costs nothing extra, which is unusual. Most providers attach higher concurrency to a spend tier.

How DeepSeek pricing has changed across all these years

Date

Milestone

Source

 

10 Sep 2026

V4.1 Flash ships and cuts Flash rates: cache miss $0.22 to $0.15 off-peak, cache hit $0.007 to $0.003, output $0.66 to $0.60. Pro unchanged

Vendor

16 Aug 2026

Flash cache-miss input moves from a flat $0.14 to $0.22 off-peak and $0.44 peak. Pro moves from $0.435 to $0.66 and $1.32

Vendor

13 Aug 2026

Peak and off-peak billing announced, off-peak set at half of peak, effective 16:00 UTC on 16 Aug 2026

Vendor

24 Apr 2026

V4-Pro and V4-Flash replace deepseek-chat and deepseek-reasoner on the API

Vendor

21 Aug 2025

New rate card announced with off-peak discounts ending 5 Sep 2025 at 16:00 UTC

Vendor

Flexprice’s Take

DeepSeek prices the cache outcome instead of the cache operation, which is honest about where the cost actually sits and unusually hard to forecast against.

DeepSeek's caching mechanics deserve credit. Writing to the cache costs nothing, there's no parameter to set, and every response reports its hit and miss token counts. Charging 2% for a hit passes the infrastructure saving straight through.

DeepSeek publishes its cache persistence rules, including the admission that hit rates aren't guaranteed. More candour than this category usually offers.

Forecasting is the hard part. Your effective input rate blends two numbers 50x apart, set by a best-effort system, then doubles inside a 35-hour weekly window. Two teams with identical token volumes can see bills an order of magnitude apart. A batch scheduler can chase the off-peak rate. An interactive product can't.

Best For

Batch and agent workloads with stable prompt prefixes and flexible scheduling.

Watch Out For

Budgeting from the headline rate, which is neither of the two you'll pay.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-model, per-customer cost tracking?

Flexprice meters it and reports margin by account.

Flexprice’s Take

DeepSeek prices the cache outcome instead of the cache operation, which is honest about where the cost actually sits and unusually hard to forecast against.

DeepSeek's caching mechanics deserve credit. Writing to the cache costs nothing, there's no parameter to set, and every response reports its hit and miss token counts. Charging 2% for a hit passes the infrastructure saving straight through.

DeepSeek publishes its cache persistence rules, including the admission that hit rates aren't guaranteed. More candour than this category usually offers.

Forecasting is the hard part. Your effective input rate blends two numbers 50x apart, set by a best-effort system, then doubles inside a 35-hour weekly window. Two teams with identical token volumes can see bills an order of magnitude apart. A batch scheduler can chase the off-peak rate. An interactive product can't.

Best For

Batch and agent workloads with stable prompt prefixes and flexible scheduling.

Watch Out For

Budgeting from the headline rate, which is neither of the two you'll pay.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-model, per-customer cost tracking?

Flexprice meters it and reports margin by account.

Customer
Sentiment Highlights

This notice is making me nervous

DeepSeek API user reacting to the announced price rise, Hacker News, August 2026

Frequently Asked Questions

Frequently Asked Questions

How much does DeepSeek cost per million tokens?

What are DeepSeek's peak hours?

Do I need to enable DeepSeek prompt caching?

Is there a free tier on the DeepSeek API?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack