Tested August 7, 2026. There is no universal winner. In our controlled sample, DeepSeek V4 Flash was faster and much cheaper on the short API task and scored higher on the semantic coding rubric; both APIs were exact on normalized document extraction, structured orders, instruction adherence, and the target-64k retrieval fixture. Claude Sonnet 5 has documented native PDF and hosted web-search API features that the checked DeepSeek Chat Completions API does not document. In separate one-attempt web-interface tests, Claude passed both the PDF and instruction cases while DeepSeek did not.
Current DeepSeek contract, rechecked August 24, 2026: the experimental deepseek-v4-flash-vision-exp model now accepts images, but DeepSeek still does not document native PDF or general-document input. The August 7 benchmark used text-only Flash before Vision Exp was released, so no score, cost, or N/A cell below has been retroactively changed.
DeepSeek vs Claude: the quick answer
| Workload | Observed result | Boundary |
|---|---|---|
| Coding API | DeepSeek 5/5; Claude 4/5 in all repeats | Semantic grader; three repeats of one debugging case |
| Normalized document API | Tie: both 3/3 exact in all three repeats | Text-normalized input, not native PDF |
| Structured output API | Tie: both 5/5 in all three valid repeats | Claude reruns used request builder v1.0.3 |
| Instruction API | Tie: both strict in 3/3 repeats | One formatting case |
| Target-64k API retrieval | Tie: both 4/4 exact at beginning, middle, and end | Same text produced 52,566-52,567 DeepSeek input tokens and 63,721 Claude input tokens |
| Short API latency | DeepSeek median total time 1.285 s; Claude 2.335 s | Five sequential requests from one run environment |
| Web citations API | Claude: 0/3 strict full passes | One dated question; manual adjudication by one editor |
| Native PDF API | Claude exact in 3/3 table and 3/3 chart repeats | DeepSeek native file input was not applicable in the checked API |
| Web UI PDF | Claude 3/3 fields; DeepSeek 1/3 | One attempt per product |
| Web UI instruction | Claude pass; DeepSeek fail | One attempt per product |
What we tested
- API balanced lane:
deepseek-v4-flashwith thinking enabled and High effort versusclaude-sonnet-5with adaptive thinking and High effort. - Quality repeats: three valid repeats per provider unless explicitly stated; five repeats for the short latency case.
- UI lane: DeepSeek Chat Instant and Claude web Sonnet 5 High, in fresh chats. UI evidence is never combined with API scores.
- Protocol: frozen English prompts, synthetic fixtures, answer keys, graders, provider order, usage records, cost rates, and exclusion rules.
- Aggregation: each attempt is scored against its rubric; a case result is the median of included repeats.
Method, exclusions, and sample limits
The main results include only successful evidence cells with valid request and fixture metadata. Invalid builder attempts were excluded before comparison: the first DeepSeek instruction attempt used an output ceiling that left no reasoning headroom; four Claude structured-output attempts used unsupported array constraints; the original long-context calibration requests were not 64K requests; and two v1.0.2 target-64k records omitted required named-tier metadata. Structured-output and target-64k results therefore use the complete v1.0.3 evidence. We did not selectively discard slow or inconvenient successful runs.
This is a small synthetic benchmark, not a population estimate. Coding, document, structured-output, instruction, and target-64k comparisons have three valid repeats per provider. Latency has five. Each UI case has one attempt per product. Exact UI completion timing, UI tokens, UI cost, regional latency, production reliability, safety, and multilingual quality were not measured.
Web UI results: PDF and instruction tests
Native PDF table result
The same synthetic one-page PDF and prompt asked for South, numeric 590, and NW-Q2-731 as JSON. Claude Sonnet 5 High returned all three fields exactly. DeepSeek Instant returned one exact field: it represented total_units as the string "590" and changed the approval code to NW- Q2- 731. This was one attempt per web product and does not rank general PDF, OCR, or chart quality.


Strict instruction result
In a separate four-line formatting case, Claude Sonnet 5 High passed every strict check. DeepSeek Instant failed: its second line contained six words and its final line was END. instead of END. No screenshot or UI latency claim is made for this case because viewport captures timed out and exact completion events were not recorded.
Balanced API quality results
| Case | DeepSeek V4 Flash | Claude Sonnet 5 | Verdict |
|---|---|---|---|
| Coding debug | 5/5 in 3/3 repeats | 4/5 in 3/3 repeats | DeepSeek led the semantic rubric. Both found and fixed the bug; Claude returned expected numeric outputs as strings. |
| Normalized document | 3/3 exact in 3/3 repeats | 3/3 exact in 3/3 repeats | Tie |
| Structured orders | 5/5 in 3/3 repeats | 5/5 in 3/3 valid v1.0.3 repeats | Tie |
| Instruction stack | Strict pass in 3/3 repeats | Strict pass in 3/3 repeats | Tie on this case |
| Target-64k retrieval | 4/4 exact in 3/3 positions | 4/4 exact in 3/3 positions | Tie on this needle task |

Latency, tokens, and measured API cost
Costs below are medians per included request, calculated from provider-reported usage and the frozen August 7, 2026 rate snapshot. They are not subscription prices. Output-token counts follow each provider’s usage accounting and include reasoning where the provider reports it inside output usage.
| Case | DeepSeek median input / output | DeepSeek median cost | Claude median input / output | Claude median cost |
|---|---|---|---|---|
| Coding | 341 / 249 | $0.00008794 | 358 / 281 | $0.003526 |
| Normalized document | 262 / 127 | $0.00003712 | 234 / 52 | $0.000988 |
| Structured orders | 398 / 687 | $0.00019540 | 512 / 149 | $0.002514 |
| Instruction stack | 175 / 253 | $0.00007778 | 140 / 363 | $0.003910 |
| Short latency | 89 / 17 | $0.00001722 | 16 / 4 | $0.000072 |
| Target-64k | 52,567 / 341 | $0.00463990 | 63,721 / 57 | $0.128012 |
On the five short sequential requests, DeepSeek’s median total time was 1.285 seconds and Claude’s was 2.335 seconds. The normalized exact-text grader passed all five attempts for both providers. This one route, region, load window, prompt, and configuration does not establish general service speed. DeepSeek had the lower measured cost in every matched case above, but lower token cost is not the same as lower total workflow cost.


Native PDF and web citations through the API
Claude’s native PDF API path returned exact table fields in 3/3 repeats and exact chart fields in 3/3 repeats. There is no matched DeepSeek native-PDF score because the checked DeepSeek Flash Chat Completions path did not document native file input; scoring an unsupported cell as a failure would mix capability and quality. Vision Exp was released after this frozen test and supports image input—not PDF, DOCX, or other general-document input—so it would require a new, separately labeled image benchmark rather than a rewritten historical result.
Claude Sonnet 5 used its hosted web-search tool for three repeats of one dated official-source question. Manual adjudication found 0/3 strict full passes. Scores were 6/10, 8/10, and 6/10: the OpenAI claim was directly supported in all three, but the DeepSeek claim was directly supported in none; one repeat used a source outside the allowed official domains, and no repeat passed the no-extra-uncited-claims check. This was a final one-editor adjudication, not the planned two-blinded-adjudicator process, so it is a limited case result rather than a general citation-quality rate. DeepSeek was not tested because a hosted search tool was not documented in the checked API.
Historical boundary: that August 7 lane used Chat Completions. DeepSeek now documents server-side web_search on the separate Responses API, which was outside the frozen protocol and has no score here.
Documented API capabilities and listed prices
| Item | DeepSeek V4 API | Claude Sonnet 5 |
|---|---|---|
| Current model status | deepseek-v4-flash: V4-Flash-0731 public beta; deepseek-v4-pro: V4-Pro-0813 GA; deepseek-v4-flash-vision-exp: experimental image model | See Anthropic’s current model page |
| Documented context / max output | 1,000,000 / 384,000 tokens for all three current IDs | 1,000,000 / 128,000 tokens |
| Image and document input | Vision Exp accepts JPEG, PNG, GIF, or WebP by URL, Base64 data URL, or image file_id. PDF and general documents remain unsupported. | Native visual PDF input documented |
| Hosted web search API | Responses: server-side web_search is documented for Flash and Pro; Chat Completions has no hosted search | Documented |
| Structured output | Chat: json_object; Responses: text.format with json_schema | Strict JSON Schema output documented |
| Responses API | All three current IDs; stateless | Provider-specific Messages API |
| Thinking controls | All three support thinking and non-thinking modes. The documented Flash/Pro Chat mapping is low → Low; medium/high/xhigh → High; max → Max. | See Anthropic’s current adaptive-thinking controls |
| Uncached input / output per 1M tokens | Time-dependent schedule below | $2.00 / $10.00 current permanent rate; historical snapshot label: “$2.00 / $10.00 introductory rate checked August 7” |
Anthropic made Claude Sonnet 5 pricing of $2 input and $10 output per 1M tokens permanent on August 10, 2026, cancelling the previously announced move to $3 / $15. The August 7 measured costs remain unchanged because they already used $2 / $10. Recheck current terms before implementation. Official sources: DeepSeek models and pricing, updates, thinking mode, Vision guide, image Files API, Chat Completions, and Responses API; Anthropic models, PDF support, structured outputs, and web search.
| DeepSeek model and effective period | Cached input | Uncached input | Output |
|---|---|---|---|
| Flash — historical, through Aug 16, 15:59 UTC | $0.0028 | $0.14 | $0.28 |
| Pro — historical, through Aug 16, 15:59 UTC | $0.003625 | $0.435 | $0.87 |
| Flash — current off-peak | $0.007 | $0.22 | $0.66 |
| Flash — current peak | $0.014 | $0.44 | $1.32 |
| Vision Exp — current off-peak | $0.007 | $0.22 | $0.66 |
| Vision Exp — current peak | $0.014 | $0.44 | $1.32 |
| Pro — current off-peak | $0.022 | $0.66 | $1.98 |
| Pro — current peak | $0.044 | $1.32 | $3.96 |
Chat product or API: decide which surface you need
A web chat is useful for interactive reading, drafting, and occasional file work. An API is the appropriate surface when software must send repeatable requests, validate results, record usage, enforce budgets, or route work automatically. The distinction is practical, not cosmetic. A chat plan may bundle native upload controls and hidden product routing, while an API exposes a named model, request body, token accounting, and documented tools. That is why the DeepSeek Instant and Claude Sonnet 5 web results above must not be used to predict the behavior of deepseek-v4-flash or claude-sonnet-5.
For a chat decision, test the exact plan, visible model, file control, search toggle, and fresh-chat behavior your users will see. For an API decision, freeze the model ID, reasoning setting, output ceiling, tool definitions, schema, retry policy, and maximum spend. Also record provider-reported input, cached-input, output, and tool usage. The same business task can favor different providers on the two surfaces, as our separate UI instruction result illustrates.
Coding workflows: a model call is not a coding agent
Our coding result measures one bounded debugging response. It does not test repository navigation, multi-file edits, shell commands, test execution, pull-request review, or long agent sessions. In the semantic grader, DeepSeek returned the fully expected structure in all three repeats; Claude found and fixed the same bug but represented expected numeric outputs as strings. That is evidence for this output contract, not proof that one system is the better coding agent.
Anthropic documents Claude Code as a local developer tool that connects to Anthropic’s API by default and can also authenticate through supported Claude plans or enterprise cloud platforms. Its official CLI reference includes interactive and print modes, JSON and streaming JSON output, turn limits, model selection, resume controls, and permission modes. Those are product-workflow capabilities around the model. They were not exercised by our API prompt.
DeepSeek documents a general Chat Completions API plus tool calling, including a separate beta strict-schema path for function calls. That makes it possible to place DeepSeek behind your own editor extension, terminal agent, test runner, or orchestration layer. The integration owner then carries more responsibility for permissions, sandboxing, file selection, command approval, diffs, retries, and audit logs. Compare the complete workflow: model accuracy, tool policy, context construction, test discipline, latency, and cost per accepted change.
Deployment and self-hosting
The hosted APIs in this benchmark are managed services. DeepSeek also maintains a verified DeepSeek V4 model collection with V4 Flash, V4 Pro, and base-weight entries. That creates a separate deployment option: evaluate the applicable model card and license, then run compatible weights on infrastructure you control or through a hosting partner. Do not assume that a self-hosted checkpoint is operationally identical to the tested hosted API. Quantization, serving engine, precision, prompt template, context implementation, hardware, and decoding settings can change output and latency.
Self-hosting can increase control over network paths, logs, model version pinning, and upgrade timing, but it also transfers capacity planning, security patching, observability, abuse controls, scaling, and incident response to the operator. Very large weights require substantial hardware and serving expertise. A lower listed API token rate may be more economical than maintaining GPUs at low utilization; local or dedicated deployment may become attractive when governance, data location, predictable utilization, or customization dominates the decision. Calculate total cost of ownership rather than comparing a token price with a GPU purchase price.
Claude Sonnet 5 is evaluated here through Anthropic’s managed API. Anthropic also documents Claude Code authentication through its own service and supported enterprise cloud platforms. This page does not claim a self-hosted Claude Sonnet 5 option, and it does not benchmark Bedrock, Vertex AI, or any third-party host. Cloud route, region, contractual terms, feature availability, and price can differ, so treat each deployment path as a new evaluation cell.
Privacy and data-control decision factors
Do not upload confidential source code, customer files, credentials, regulated records, or private identifiers merely because a benchmark fixture was safe. Our fixtures were synthetic. Before production, document the exact product and account type, processing region, retention period, training setting, cache behavior, subprocessors, deletion controls, incident process, and whether web or file tools send data to additional systems.
Anthropic’s official commercial-product privacy guidance says standard API inputs and outputs are deleted from its backend within 30 days, subject to stated exceptions and different agreements. Its separate zero-data-retention guidance says approved arrangements apply to the API and products using the commercial organization API key, while some features—including Files API persistence, explicit caching, certain batch operations, beta products, or third-party web search—can have different handling. Consumer Claude and Claude Code sessions authenticated with consumer plans have a separate training-choice policy. Verify the current terms for your account; do not transfer an API rule to a consumer chat plan.
DeepSeek’s API documentation describes per-user_id content-safety, scheduling, and KV-cache isolation and explicitly warns not to put private information in that identifier. Its caching documentation says cache entries are isolated between users and unused entries are cleared after a period. Those statements are useful controls, but they are not a complete security assessment or a substitute for the current contract and privacy policy. If data residency or retention is decisive, obtain written terms for the exact service or evaluate an eligible self-managed weight deployment.
- Redact secrets and minimize files before any model call.
- Keep API keys in a secret manager and out of prompts, repositories, screenshots, and browser code.
- Use least-privilege tools, explicit command approval, isolated execution, output validation, and human review proportional to harm.
- Log model IDs and configuration without logging sensitive prompt bodies unnecessarily.
- Test deletion, retention, and access-control procedures before processing production data.
Which should you choose?
- Choose DeepSeek for cost-sensitive text API work, or evaluate Vision Exp when the input is an image rather than a general document. Normalize PDFs and other unsupported documents yourself, and validate every output. DeepSeek was cheaper across the dated matched requests and faster on the short latency case.
- Choose Claude for native PDF or hosted-search API workflows when those documented provider features reduce your engineering burden. Still validate citations: this small search case had no strict full pass.
- Test both for code, long context, images, and exact automation. The observed winners and ties are case-specific. Route workloads rather than forcing one global model choice.
For implementation details, see the DeepSeek API guide, DeepSeek pricing guide, API cost calculator, context and output limits guide, and the comparison hub. The independent DeepSeek guide covers the broader product map.
Reproducibility record
- Snapshot: August 7, 2026.
- API models:
deepseek-v4-flashandclaude-sonnet-5, balanced High-effort settings. - Builder: v1.0.3 for corrected structured-output and target-64k evidence; other main cases use evidence selected by the final aggregator.
- Target-64k tier:
target-64k-v2, three positions, not the excluded calibration fixture. - PDF SHA-256:
cb3259211734feda0265ecf46cc76745e116d72782b01e5e3c136d9fa7b0a07f. - Public dataset: none. Dataset schema and download claims remain disabled until a versioned public release exists.
Limitations
- Small English synthetic cases do not represent all real workloads, languages, layouts, regions, or future model versions.
- The UI labels are not proof of underlying API model IDs; UI and API results are separate.
- The citation result uses one question and one final editor; no inter-rater agreement was recorded.
- Latency reflects one sequential run environment and does not estimate tail latency or uptime.
- Costs exclude subscriptions, taxes, engineering, retries, and future price changes.
- This site is independent and is not affiliated with DeepSeek or Anthropic.
Frequently asked questions
Is DeepSeek better than Claude?
Not universally. DeepSeek led this coding rubric, short latency, and measured API cost. Claude offers documented native PDF and hosted-search API paths and led both one-attempt UI cases.
Is DeepSeek cheaper than Claude?
Yes for every matched API request in this dated sample and in the listed token rates, but total workflow cost also depends on validation, retries, tools, and engineering.
Which is better for PDF analysis?
Claude documents native visual PDF input and was exact in our Claude-only API PDF cases. In the separate one-attempt UI table case, Claude returned 3/3 exact fields and DeepSeek returned 1/3. These results do not rank all PDFs or OCR tasks.
What happened in the target-64k test?
Both APIs returned all four exact fields at beginning, middle, and end positions. The same fixture counted as about 52.6K DeepSeek input tokens and 63.7K Claude input tokens, so it is a named target tier rather than a claim of identical provider token counts.
Did Claude pass the web-citation test?
No repeat achieved a strict full pass: 0/3. The result is limited to one dated question and a one-editor manual adjudication.
Are web app results the same as API results?
No. Plans, visible modes, native tools, hidden routing, limits, and fallbacks can differ. This page keeps the two surfaces separate.
Which is better for coding, DeepSeek or Claude?
DeepSeek led our one semantic debugging rubric, while both providers found and fixed the bug. That test did not compare full coding agents. Evaluate repository navigation, edit quality, tests, permission controls, latency, and accepted-change cost in your own environment.
Can I self-host DeepSeek or Claude?
DeepSeek publishes official V4 model entries that can support a self-managed evaluation, subject to the applicable model card, license, hardware, and serving stack. Claude Sonnet 5 is compared here as a managed service; this page does not claim downloadable Claude weights. Self-hosted and hosted results are not interchangeable.
Which is more private, DeepSeek or Claude?
There is no safe answer without naming the product, account, deployment, region, tools, retention agreement, and data. Compare current written terms and controls for the exact route, use synthetic tests, minimize data, and seek security or legal review for regulated workloads.
