DeepSeek vs Claude: API and Web App Tests (2026)

Tested August 7, 2026. There is no universal winner. In our controlled sample, DeepSeek V4 Flash was faster and much cheaper on the short API task and scored higher on the semantic coding rubric; both APIs were exact on normalized document extraction, structured orders, instruction adherence, and the target-64k retrieval fixture. Claude Sonnet 5 has documented native PDF and hosted web-search API features that the checked DeepSeek Chat Completions API does not document. In separate one-attempt web-interface tests, Claude passed both the PDF and instruction cases while DeepSeek did not.

Current DeepSeek contract, rechecked August 24, 2026: the experimental deepseek-v4-flash-vision-exp model now accepts images, but DeepSeek still does not document native PDF or general-document input. The August 7 benchmark used text-only Flash before Vision Exp was released, so no score, cost, or N/A cell below has been retroactively changed.

DeepSeek vs Claude: the quick answer

WorkloadObserved resultBoundary
Coding APIDeepSeek 5/5; Claude 4/5 in all repeatsSemantic grader; three repeats of one debugging case
Normalized document APITie: both 3/3 exact in all three repeatsText-normalized input, not native PDF
Structured output APITie: both 5/5 in all three valid repeatsClaude reruns used request builder v1.0.3
Instruction APITie: both strict in 3/3 repeatsOne formatting case
Target-64k API retrievalTie: both 4/4 exact at beginning, middle, and endSame text produced 52,566-52,567 DeepSeek input tokens and 63,721 Claude input tokens
Short API latencyDeepSeek median total time 1.285 s; Claude 2.335 sFive sequential requests from one run environment
Web citations APIClaude: 0/3 strict full passesOne dated question; manual adjudication by one editor
Native PDF APIClaude exact in 3/3 table and 3/3 chart repeatsDeepSeek native file input was not applicable in the checked API
Web UI PDFClaude 3/3 fields; DeepSeek 1/3One attempt per product
Web UI instructionClaude pass; DeepSeek failOne attempt per product

What we tested

  • API balanced lane: deepseek-v4-flash with thinking enabled and High effort versus claude-sonnet-5 with adaptive thinking and High effort.
  • Quality repeats: three valid repeats per provider unless explicitly stated; five repeats for the short latency case.
  • UI lane: DeepSeek Chat Instant and Claude web Sonnet 5 High, in fresh chats. UI evidence is never combined with API scores.
  • Protocol: frozen English prompts, synthetic fixtures, answer keys, graders, provider order, usage records, cost rates, and exclusion rules.
  • Aggregation: each attempt is scored against its rubric; a case result is the median of included repeats.

Method, exclusions, and sample limits

The main results include only successful evidence cells with valid request and fixture metadata. Invalid builder attempts were excluded before comparison: the first DeepSeek instruction attempt used an output ceiling that left no reasoning headroom; four Claude structured-output attempts used unsupported array constraints; the original long-context calibration requests were not 64K requests; and two v1.0.2 target-64k records omitted required named-tier metadata. Structured-output and target-64k results therefore use the complete v1.0.3 evidence. We did not selectively discard slow or inconvenient successful runs.

This is a small synthetic benchmark, not a population estimate. Coding, document, structured-output, instruction, and target-64k comparisons have three valid repeats per provider. Latency has five. Each UI case has one attempt per product. Exact UI completion timing, UI tokens, UI cost, regional latency, production reliability, safety, and multilingual quality were not measured.

Web UI results: PDF and instruction tests

Native PDF table result

The same synthetic one-page PDF and prompt asked for South, numeric 590, and NW-Q2-731 as JSON. Claude Sonnet 5 High returned all three fields exactly. DeepSeek Instant returned one exact field: it represented total_units as the string "590" and changed the approval code to NW- Q2- 731. This was one attempt per web product and does not rank general PDF, OCR, or chart quality.

Claude Sonnet 5 returning all three exact values from the synthetic PDF table test
Claude Sonnet 5 High returned all three answer-key fields exactly in the controlled web-interface PDF case.
DeepSeek Instant returning one exact field, a string total, and a spaced approval code in the synthetic PDF table test
DeepSeek Instant returned the region exactly, represented total_units as a string, and inserted spaces into the approval code.

Strict instruction result

In a separate four-line formatting case, Claude Sonnet 5 High passed every strict check. DeepSeek Instant failed: its second line contained six words and its final line was END. instead of END. No screenshot or UI latency claim is made for this case because viewport captures timed out and exact completion events were not recorded.

Balanced API quality results

CaseDeepSeek V4 FlashClaude Sonnet 5Verdict
Coding debug5/5 in 3/3 repeats4/5 in 3/3 repeatsDeepSeek led the semantic rubric. Both found and fixed the bug; Claude returned expected numeric outputs as strings.
Normalized document3/3 exact in 3/3 repeats3/3 exact in 3/3 repeatsTie
Structured orders5/5 in 3/3 repeats5/5 in 3/3 valid v1.0.3 repeatsTie
Instruction stackStrict pass in 3/3 repeatsStrict pass in 3/3 repeatsTie on this case
Target-64k retrieval4/4 exact in 3/3 positions4/4 exact in 3/3 positionsTie on this needle task
Grouped bars compare strict API pass rates across nine tasks. DeepSeek passed all six applicable tasks and had three N/A cells; Claude passed seven tasks and failed the strict code-diagnosis and citation gates.
Strict or exact pass rate for the completed balanced-lane API tasks. N/A marks a task that was not applicable in the checked API protocol and is not scored as a failure. Tests run August 7, 2026.

Latency, tokens, and measured API cost

Costs below are medians per included request, calculated from provider-reported usage and the frozen August 7, 2026 rate snapshot. They are not subscription prices. Output-token counts follow each provider’s usage accounting and include reasoning where the provider reports it inside output usage.

CaseDeepSeek median input / outputDeepSeek median costClaude median input / outputClaude median cost
Coding341 / 249$0.00008794358 / 281$0.003526
Normalized document262 / 127$0.00003712234 / 52$0.000988
Structured orders398 / 687$0.00019540512 / 149$0.002514
Instruction stack175 / 253$0.00007778140 / 363$0.003910
Short latency89 / 17$0.0000172216 / 4$0.000072
Target-64k52,567 / 341$0.0046399063,721 / 57$0.128012

On the five short sequential requests, DeepSeek’s median total time was 1.285 seconds and Claude’s was 2.335 seconds. The normalized exact-text grader passed all five attempts for both providers. This one route, region, load window, prompt, and configuration does not establish general service speed. DeepSeek had the lower measured cost in every matched case above, but lower token cost is not the same as lower total workflow cost.

Five short-response TTFT runs. DeepSeek median 1.284 seconds; Claude median 2.320 seconds. Individual run values are shown as dots.
Time to first visible text for five sequential streaming runs per provider using the same exact READY prompt, concurrency one, on August 7, 2026.
Target-64K same-text retrieval test. Both providers passed 3 of 3 positions. DeepSeek reported 52,566 to 52,567 input tokens, median total time 4.268 seconds, and median cost $0.004640. Claude reported 63,721 input tokens, median total time 4.157 seconds, and median cost $0.128012.
Matched 89,610-byte target-64K retrieval test at beginning, middle, and end positions. Actual provider-reported input-token counts, total response time, and per-run API cost are shown. Tests run August 7, 2026.

Native PDF and web citations through the API

Claude’s native PDF API path returned exact table fields in 3/3 repeats and exact chart fields in 3/3 repeats. There is no matched DeepSeek native-PDF score because the checked DeepSeek Flash Chat Completions path did not document native file input; scoring an unsupported cell as a failure would mix capability and quality. Vision Exp was released after this frozen test and supports image input—not PDF, DOCX, or other general-document input—so it would require a new, separately labeled image benchmark rather than a rewritten historical result.

Claude Sonnet 5 used its hosted web-search tool for three repeats of one dated official-source question. Manual adjudication found 0/3 strict full passes. Scores were 6/10, 8/10, and 6/10: the OpenAI claim was directly supported in all three, but the DeepSeek claim was directly supported in none; one repeat used a source outside the allowed official domains, and no repeat passed the no-extra-uncited-claims check. This was a final one-editor adjudication, not the planned two-blinded-adjudicator process, so it is a limited case result rather than a general citation-quality rate. DeepSeek was not tested because a hosted search tool was not documented in the checked API.

Historical boundary: that August 7 lane used Chat Completions. DeepSeek now documents server-side web_search on the separate Responses API, which was outside the frozen protocol and has no score here.

Documented API capabilities and listed prices

ItemDeepSeek V4 APIClaude Sonnet 5
Current model statusdeepseek-v4-flash: V4-Flash-0731 public beta; deepseek-v4-pro: V4-Pro-0813 GA; deepseek-v4-flash-vision-exp: experimental image modelSee Anthropic’s current model page
Documented context / max output1,000,000 / 384,000 tokens for all three current IDs1,000,000 / 128,000 tokens
Image and document inputVision Exp accepts JPEG, PNG, GIF, or WebP by URL, Base64 data URL, or image file_id. PDF and general documents remain unsupported.Native visual PDF input documented
Hosted web search APIResponses: server-side web_search is documented for Flash and Pro; Chat Completions has no hosted searchDocumented
Structured outputChat: json_object; Responses: text.format with json_schemaStrict JSON Schema output documented
Responses APIAll three current IDs; statelessProvider-specific Messages API
Thinking controlsAll three support thinking and non-thinking modes. The documented Flash/Pro Chat mapping is low → Low; medium/high/xhigh → High; max → Max.See Anthropic’s current adaptive-thinking controls
Uncached input / output per 1M tokensTime-dependent schedule below$2.00 / $10.00 current permanent rate; historical snapshot label: “$2.00 / $10.00 introductory rate checked August 7”
Current DeepSeek contract and pricing documentation rechecked August 24, 2026; Claude facts retain their separately stated verification dates, and all August 7 test evidence remains frozen.

Anthropic made Claude Sonnet 5 pricing of $2 input and $10 output per 1M tokens permanent on August 10, 2026, cancelling the previously announced move to $3 / $15. The August 7 measured costs remain unchanged because they already used $2 / $10. Recheck current terms before implementation. Official sources: DeepSeek models and pricing, updates, thinking mode, Vision guide, image Files API, Chat Completions, and Responses API; Anthropic models, PDF support, structured outputs, and web search.

DeepSeek model and effective periodCached inputUncached inputOutput
Flash — historical, through Aug 16, 15:59 UTC$0.0028$0.14$0.28
Pro — historical, through Aug 16, 15:59 UTC$0.003625$0.435$0.87
Flash — current off-peak$0.007$0.22$0.66
Flash — current peak$0.014$0.44$1.32
Vision Exp — current off-peak$0.007$0.22$0.66
Vision Exp — current peak$0.014$0.44$1.32
Pro — current off-peak$0.022$0.66$1.98
Pro — current peak$0.044$1.32$3.96
USD per 1M tokens. DeepSeek peak windows are 01:00–04:00 and 06:00–10:00 UTC from Monday through Friday. All other weekday hours are off-peak, and Saturday and Sunday are off-peak all day. Vision Exp shares Flash token rates; image inputs are converted to billed input tokens. The measured August 7 costs above remain unchanged.

Chat product or API: decide which surface you need

A web chat is useful for interactive reading, drafting, and occasional file work. An API is the appropriate surface when software must send repeatable requests, validate results, record usage, enforce budgets, or route work automatically. The distinction is practical, not cosmetic. A chat plan may bundle native upload controls and hidden product routing, while an API exposes a named model, request body, token accounting, and documented tools. That is why the DeepSeek Instant and Claude Sonnet 5 web results above must not be used to predict the behavior of deepseek-v4-flash or claude-sonnet-5.

For a chat decision, test the exact plan, visible model, file control, search toggle, and fresh-chat behavior your users will see. For an API decision, freeze the model ID, reasoning setting, output ceiling, tool definitions, schema, retry policy, and maximum spend. Also record provider-reported input, cached-input, output, and tool usage. The same business task can favor different providers on the two surfaces, as our separate UI instruction result illustrates.

Coding workflows: a model call is not a coding agent

Our coding result measures one bounded debugging response. It does not test repository navigation, multi-file edits, shell commands, test execution, pull-request review, or long agent sessions. In the semantic grader, DeepSeek returned the fully expected structure in all three repeats; Claude found and fixed the same bug but represented expected numeric outputs as strings. That is evidence for this output contract, not proof that one system is the better coding agent.

Anthropic documents Claude Code as a local developer tool that connects to Anthropic’s API by default and can also authenticate through supported Claude plans or enterprise cloud platforms. Its official CLI reference includes interactive and print modes, JSON and streaming JSON output, turn limits, model selection, resume controls, and permission modes. Those are product-workflow capabilities around the model. They were not exercised by our API prompt.

DeepSeek documents a general Chat Completions API plus tool calling, including a separate beta strict-schema path for function calls. That makes it possible to place DeepSeek behind your own editor extension, terminal agent, test runner, or orchestration layer. The integration owner then carries more responsibility for permissions, sandboxing, file selection, command approval, diffs, retries, and audit logs. Compare the complete workflow: model accuracy, tool policy, context construction, test discipline, latency, and cost per accepted change.

Deployment and self-hosting

The hosted APIs in this benchmark are managed services. DeepSeek also maintains a verified DeepSeek V4 model collection with V4 Flash, V4 Pro, and base-weight entries. That creates a separate deployment option: evaluate the applicable model card and license, then run compatible weights on infrastructure you control or through a hosting partner. Do not assume that a self-hosted checkpoint is operationally identical to the tested hosted API. Quantization, serving engine, precision, prompt template, context implementation, hardware, and decoding settings can change output and latency.

Self-hosting can increase control over network paths, logs, model version pinning, and upgrade timing, but it also transfers capacity planning, security patching, observability, abuse controls, scaling, and incident response to the operator. Very large weights require substantial hardware and serving expertise. A lower listed API token rate may be more economical than maintaining GPUs at low utilization; local or dedicated deployment may become attractive when governance, data location, predictable utilization, or customization dominates the decision. Calculate total cost of ownership rather than comparing a token price with a GPU purchase price.

Claude Sonnet 5 is evaluated here through Anthropic’s managed API. Anthropic also documents Claude Code authentication through its own service and supported enterprise cloud platforms. This page does not claim a self-hosted Claude Sonnet 5 option, and it does not benchmark Bedrock, Vertex AI, or any third-party host. Cloud route, region, contractual terms, feature availability, and price can differ, so treat each deployment path as a new evaluation cell.

Privacy and data-control decision factors

Do not upload confidential source code, customer files, credentials, regulated records, or private identifiers merely because a benchmark fixture was safe. Our fixtures were synthetic. Before production, document the exact product and account type, processing region, retention period, training setting, cache behavior, subprocessors, deletion controls, incident process, and whether web or file tools send data to additional systems.

Anthropic’s official commercial-product privacy guidance says standard API inputs and outputs are deleted from its backend within 30 days, subject to stated exceptions and different agreements. Its separate zero-data-retention guidance says approved arrangements apply to the API and products using the commercial organization API key, while some features—including Files API persistence, explicit caching, certain batch operations, beta products, or third-party web search—can have different handling. Consumer Claude and Claude Code sessions authenticated with consumer plans have a separate training-choice policy. Verify the current terms for your account; do not transfer an API rule to a consumer chat plan.

DeepSeek’s API documentation describes per-user_id content-safety, scheduling, and KV-cache isolation and explicitly warns not to put private information in that identifier. Its caching documentation says cache entries are isolated between users and unused entries are cleared after a period. Those statements are useful controls, but they are not a complete security assessment or a substitute for the current contract and privacy policy. If data residency or retention is decisive, obtain written terms for the exact service or evaluate an eligible self-managed weight deployment.

  • Redact secrets and minimize files before any model call.
  • Keep API keys in a secret manager and out of prompts, repositories, screenshots, and browser code.
  • Use least-privilege tools, explicit command approval, isolated execution, output validation, and human review proportional to harm.
  • Log model IDs and configuration without logging sensitive prompt bodies unnecessarily.
  • Test deletion, retention, and access-control procedures before processing production data.

Which should you choose?

  • Choose DeepSeek for cost-sensitive text API work, or evaluate Vision Exp when the input is an image rather than a general document. Normalize PDFs and other unsupported documents yourself, and validate every output. DeepSeek was cheaper across the dated matched requests and faster on the short latency case.
  • Choose Claude for native PDF or hosted-search API workflows when those documented provider features reduce your engineering burden. Still validate citations: this small search case had no strict full pass.
  • Test both for code, long context, images, and exact automation. The observed winners and ties are case-specific. Route workloads rather than forcing one global model choice.

For implementation details, see the DeepSeek API guide, DeepSeek pricing guide, API cost calculator, context and output limits guide, and the comparison hub. The independent DeepSeek guide covers the broader product map.

Reproducibility record

  • Snapshot: August 7, 2026.
  • API models: deepseek-v4-flash and claude-sonnet-5, balanced High-effort settings.
  • Builder: v1.0.3 for corrected structured-output and target-64k evidence; other main cases use evidence selected by the final aggregator.
  • Target-64k tier: target-64k-v2, three positions, not the excluded calibration fixture.
  • PDF SHA-256: cb3259211734feda0265ecf46cc76745e116d72782b01e5e3c136d9fa7b0a07f.
  • Public dataset: none. Dataset schema and download claims remain disabled until a versioned public release exists.

Limitations

  • Small English synthetic cases do not represent all real workloads, languages, layouts, regions, or future model versions.
  • The UI labels are not proof of underlying API model IDs; UI and API results are separate.
  • The citation result uses one question and one final editor; no inter-rater agreement was recorded.
  • Latency reflects one sequential run environment and does not estimate tail latency or uptime.
  • Costs exclude subscriptions, taxes, engineering, retries, and future price changes.
  • This site is independent and is not affiliated with DeepSeek or Anthropic.

Frequently asked questions

Is DeepSeek better than Claude?

Not universally. DeepSeek led this coding rubric, short latency, and measured API cost. Claude offers documented native PDF and hosted-search API paths and led both one-attempt UI cases.

Is DeepSeek cheaper than Claude?

Yes for every matched API request in this dated sample and in the listed token rates, but total workflow cost also depends on validation, retries, tools, and engineering.

Which is better for PDF analysis?

Claude documents native visual PDF input and was exact in our Claude-only API PDF cases. In the separate one-attempt UI table case, Claude returned 3/3 exact fields and DeepSeek returned 1/3. These results do not rank all PDFs or OCR tasks.

What happened in the target-64k test?

Both APIs returned all four exact fields at beginning, middle, and end positions. The same fixture counted as about 52.6K DeepSeek input tokens and 63.7K Claude input tokens, so it is a named target tier rather than a claim of identical provider token counts.

Did Claude pass the web-citation test?

No repeat achieved a strict full pass: 0/3. The result is limited to one dated question and a one-editor manual adjudication.

Are web app results the same as API results?

No. Plans, visible modes, native tools, hidden routing, limits, and fallbacks can differ. This page keeps the two surfaces separate.

Which is better for coding, DeepSeek or Claude?

DeepSeek led our one semantic debugging rubric, while both providers found and fixed the bug. That test did not compare full coding agents. Evaluate repository navigation, edit quality, tests, permission controls, latency, and accepted-change cost in your own environment.

Can I self-host DeepSeek or Claude?

DeepSeek publishes official V4 model entries that can support a self-managed evaluation, subject to the applicable model card, license, hardware, and serving stack. Claude Sonnet 5 is compared here as a managed service; this page does not claim downloadable Claude weights. Self-hosted and hosted results are not interchangeable.

Which is more private, DeepSeek or Claude?

There is no safe answer without naming the product, account, deployment, region, tools, retention agreement, and data. Compare current written terms and controls for the exact route, use synthetic tests, minimize data, and seek security or legal review for regulated workloads.