DeepSeek Pricing: API Rates and Usage Costs

DeepSeek’s official mobile assistant is free to use, and its web chat has no advertised consumer subscription. The API is a separate service: you pay for the content the model processes and generates. Your bill depends on the model, input and output volume, cached content, and when the requests run.

This page explains direct DeepSeek API charges and the related cost of using its Harness desktop client. If you only want ordinary chat, our guide to what is free explains that route. Chat-Deep is an independent guide; the prices below belong to DeepSeek’s service.

Jump to API rates, cost examples, image costs, or payment and balance.

Current DeepSeek API prices

All prices are in US dollars per million tokens, checked against DeepSeek’s official pricing on October 4, 2026. Flash currently means DeepSeek-V4.1-Flash; Pro means DeepSeek-V4-Pro-0813. Another provider’s rates may differ.

ModelPeriodCached inputUncached inputOutput
Flash (deepseek-flash)Off-peak$0.003$0.15$0.60
FlashPeak$0.006$0.30$1.20
V4 Pro (deepseek-v4-pro)Off-peak$0.022$0.66$1.98
V4 ProPeak$0.044$1.32$3.96

On a phone, swipe the table sideways to see all the rates.

A token is a unit of content, not necessarily a whole word. Input includes your prompt and any conversation history you send. Output is what the model generates. Cached input is the portion reused from a matching earlier request; the rest is charged at the uncached rate. The official token guide explains why word counts are only a rough way to estimate usage.

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays. All other times are off-peak, including weekends and those holidays in full. Off-peak rates are half the peak rates. Convert the UTC windows to your local time before scheduling work.

Flash is the cheaper starting point if cost is your main concern. Paying more for Pro makes sense only when its results justify the difference for your task. The name “Pro” here refers to an API model, not a monthly chat membership.

What would 1,000 requests cost?

Suppose each request uses 10,000 uncached input tokens and 5,000 output tokens. At Flash’s off-peak rates, the input costs $0.0015 and the output costs $0.003. That makes $0.0045 per request, or $4.50 for 1,000 requests. At peak rates, the same workload costs $9.

Using the same assumed token counts with V4 Pro would cost $16.50 for 1,000 off-peak requests, or $33 at peak rates. This compares the bill for a fixed workload, not the models’ answer quality or the number of tokens each would actually generate.

These are calculations, not measured bills. For your own estimate, multiply each token count by its rate per million and divide by 1,000,000. Add the cached input, uncached input and output charges. If usage spans both pricing periods, calculate each portion at its applicable rate.

The same workload with some cached input

Now suppose the 1,000 Flash requests still use 10 million input tokens and 5 million output tokens in total, but the API reports 8 million cached input tokens and 2 million uncached input tokens. At off-peak rates:

  • Cached input: 8 × $0.003 = $0.024.
  • Uncached input: 2 × $0.15 = $0.30.
  • Output: 5 × $0.60 = $3.00.

The total is $3.324, or about $3.32, instead of $4.50. The large discount on cached input does not apply to the output bill. Here, output accounts for most of the cost. For your own estimate, use the cache-hit amount reported for your requests; repeating a prompt does not guarantee a particular cache-hit amount.

For a coding assistant or another app that calls the model repeatedly, one user task can involve several API requests. Budget for the entire sequence, including any conversation history sent again, rather than treating each user message as a fixed-price action.

How much do images cost in Flash?

Flash accepts images alongside text. Its vision documentation says images are converted into input tokens and billed with the text you send. The charge is based on the processed image dimensions, not simply the file’s size in megabytes.

DeepSeek resizes images before processing them. The current guide sets an upper bound of 1,024 tokens per image; smaller images do not necessarily reach that limit. Each image in a multi-image request is counted separately. A very large image can therefore use the same number of tokens as another large image after resizing.

For illustration, 1,000 images billed at 1,024 tokens each would add 1.024 million input tokens. At Flash’s off-peak uncached-input rate, that image-input portion costs $0.1536. Text instructions, conversation history and generated output cost extra. This is a calculation at the stated image-token limit, not a fixed price for 1,000 complete image-analysis tasks.

Use the official image-token estimator when planning, then reconcile against the API’s returned usage. If your total input count already includes the images, do not add the estimated image tokens again.

What changes the bill most?

Long answers and thinking use output tokens. In thinking mode, generated reasoning also counts toward output usage. A short final answer can therefore cost more than its visible length suggests. Use the API’s reported usage rather than counting the words on screen.

Repeated input can cost less, but the discount is not guaranteed. DeepSeek’s automatic context caching requires an exact match to a prefix that has already been saved in the cache. Stable instructions or repeated document content may benefit; merely similar text does not establish a hit. Cache creation takes time and cached content can be cleared. For an initial budget, use uncached rates until your usage shows how much is being reused.

In Chat Completions usage, prompt_cache_hit_tokens and prompt_cache_miss_tokens are the two parts of prompt_tokens. Charge those parts at their respective rates, rather than charging the whole input at the uncached rate and adding the cache charge again. Likewise, reasoning_tokens is a breakdown within completion_tokens, not extra output to add a second time.

Start with Flash, send the context the task needs, and request an answer of an appropriate length. Schedule work that can wait outside peak hours. Check that the results remain useful: a shorter but incomplete answer or repeated failed attempts can defeat the intended saving.

How do you pay?

API charges are deducted from your DeepSeek platform balance as you use the service. You top up rather than buy a monthly chat subscription. If your account has granted credit, it is used before the paid balance. Do not assume every new account includes a free allocation.

The official top-up help lists PayPal, bank card, Alipay and WeChat Pay. Use the platform’s Top Up page to see the options available to you, and Billing to check the transaction. Base the amount on your expected usage rather than an assumed unlimited allowance.

If a request returns 402 Insufficient Balance, check the available funds and top up as needed. The error-code guide separates this from authentication and rate-limit errors; adding money is not the fix for every failed request.

For your first integration, our DeepSeek API guide covers the key and first request. Paying for API usage does not remove another app’s separate subscription or service fees.

What about DeepSeek Harness?

Harness is a separate desktop client. Under the Harness terms, its default official-model service requires a DeepSeek account and prepaid Open Platform funds. The client being available to download does not include unrestricted free model usage.

If you configure another model provider, that provider’s charges apply. A DeepSeek top-up does not pay a different provider’s bill. For multi-step work, estimate the calls and tokens used across the task, including follow-up requests, rather than expecting one flat charge for the finished document or coding job.

Questions about your balance

Does paid DeepSeek API credit expire?

DeepSeek’s Help Center says top-up balances do not expire. That answer is about money you have added. Check the stated conditions of any granted or promotional credit separately; do not assume it has the same terms.

Can you get a refund for unused credit?

The refund help page directs online-payment users to Billing → Refunds; corporate bank-transfer users must submit a refund ticket. The Open Platform terms make refunds subject to review and any necessary fees. They cover the remaining unused balance together, not a partial amount, and exclude money already spent on usage.

Privacy and cookie settings