DeepSeek vs Qwen is not one comparison. It can mean DeepSeek Chat versus Qwen Chat, DeepSeek’s API versus Alibaba Cloud Model Studio, or DeepSeek V4 weights versus Qwen’s open-weight models. A useful answer must keep those product layers separate.
Current documentation rechecked August 16, 2026; original account tests dated July 28. DeepSeek now serves V4-Flash-0731 as a public beta and V4-Pro-0813 at GA. Alibaba Cloud’s public Model Studio catalog documents current Qwen models by region and protocol; the separate Singapore account observation below remains a dated account result, not the source of truth for global catalog availability.
We ran the same synthetic English decision fixture through DeepSeek Chat, the DeepSeek API, Qwen Studio, and five authenticated Alibaba Cloud Model Studio API configurations. Every result below is labeled as a single-run observation, not a general speed ranking. The fixture, expected answer, mode, token usage, elapsed time, and list-price estimate are published so the result can be audited.
Quick Verdict: DeepSeek vs Qwen
Choose DeepSeek first for very low cached-input pricing, a straightforward first-party API, 1M context on both V4 models, and MIT-licensed V4 weights. V4-Flash was also the lowest-cost successful API run on our fixture because it used far fewer output tokens; V4-Pro is the higher-reasoning tier.
Choose Qwen first for a broader hosted model platform, image and video input, Alibaba Cloud integration, or open-weight models that are more realistic for local and custom deployment. Qwen3.7 Plus is the documented multimodal hosted option, while Qwen3.6-27B and 35B-A3B are important current open-weight candidates.
For coding and reasoning quality, test both. Vendor benchmark tables use different harnesses, tools, token budgets, and model snapshots. A model that leads one leaderboard can still lose on your repository, schema, retry rate, or cost per accepted result.
| Decision | Better starting point | Reason | Evidence level |
|---|---|---|---|
| Lowest short-prompt list rates | Depends | Qwen3.7 Flash is lower for uncached short input/output; DeepSeek V4-Flash is lower for cached input | Official prices checked July 28 |
| Difficult coding or reasoning | Test V4-Pro vs current Qwen3.8 Max; July 28 used Qwen3.7 Max | Both can be correct, but mode and output budget materially changed our fixture; that dated Qwen3.7 result is not a Qwen3.8 result | Matched single-fixture test; broader suite required. Current Qwen3.8 specifications are a separate evidence layer. |
| Image, video, or visual documents | Qwen3.7 Plus or Flash | Both accept text, images, and video | Documented feature fit |
| Local deployment on moderate infrastructure | Qwen3.6 open weights | 27B and 35B-A3B are much smaller than DeepSeek V4 weights | Official model cards; hardware test required |
| 1M hosted context | Both | Both providers document 1M options | Maximum size only; retrieval quality untested here |
| Standard permissive license | Both | DeepSeek V4 lists MIT; current Qwen3.6 open weights list Apache 2.0 | Official repository licenses |
What Is Being Compared?
Consumer chat products
DeepSeek Chat and Qwen Studio (formerly Qwen Chat) are consumer-facing products with their own interfaces, limits, file tools, search features, accounts, and privacy terms. A free chat result does not establish API cost or the behavior of downloadable weights. Product features can also change without changing the underlying model name.
Hosted developer APIs
DeepSeek provides Chat Completions, stateless Responses, and Anthropic-compatible routes. Alibaba Cloud Model Studio provides Qwen and other models through region-specific OpenAI-compatible, Anthropic-compatible, and DashScope interfaces. Its public model catalog lists model IDs, protocols, and region-specific base URLs, including Beijing, Singapore, Tokyo, Frankfurt, and US (Virginia) for supported families. Availability still varies by exact model and deployment scope, so select the region first and use that public catalog instead of inferring global access from one account.
Open-weight models
DeepSeek V4 and selected Qwen models publish downloadable weights. That does not make the hosted aliases and downloaded checkpoints interchangeable. A provider may apply a different snapshot, prompt template, quantization, safety layer, or serving configuration.
| Layer | DeepSeek | Qwen | Use this layer for |
|---|---|---|---|
| Consumer chat | DeepSeek Chat | Qwen Studio (formerly Qwen Chat) | Interface, search, files, free access, and everyday use |
| Hosted frontier API | deepseek-v4-pro | qwen3.8-max | Hard reasoning, coding, and agents |
| Lower-cost hosted API | deepseek-v4-flash | qwen3.7-flash | Volume, latency, and cost testing |
| Hosted multimodal | Current V4 comparison is text-focused | qwen3.7-plus, qwen3.7-flash, or the Max June 8 snapshot | Images, video, and visual documents |
| Open weights | V4-Flash / V4-Pro | Qwen3.6-27B / Qwen3.6-35B-A3B | Self-hosting, tuning, and infrastructure control |
Current Models and Documented Capabilities
DeepSeek V4
DeepSeek’s official models page lists deepseek-v4-flash (V4-Flash-0731 public beta) and deepseek-v4-pro (V4-Pro-0813 GA). Both list a 1M context window, maximum output up to 384K, thinking and non-thinking modes, Chat Completions, and the stateless Responses API. Chat uses json_object and has no hosted search tool; Responses adds json_schema and server-side web_search. Reasoning low maps to Low; medium, high, and xhigh map to High; and max maps to Max.
Current Qwen hosted models and the July 28 Qwen3.7 snapshot
Alibaba’s official catalog checked August 16 identifies qwen3.8-max as its current strongest-reasoning Qwen model and lists a 1M-token context window. At the July 28 test snapshot, Singapore’s pay-as-you-go text lineup included qwen3.7-max, qwen3.7-plus, and qwen3.7-flash; Qwen3.7 Flash had launched for Singapore on July 25. The authenticated GET /models result below remains evidence only for that Singapore workspace on July 28 and is not a current Qwen3.8 inventory.
Alibaba’s current capability table marks a 1M context window, thinking, function calling, built-in tools, and structured output as supported for qwen3.8-max. Test the exact protocol and region you intend to deploy. The Qwen3.7 Max, Plus, and Flash specifications and account evidence below remain a July 28 snapshot.
Model IDs matter. Current workloads should evaluate qwen3.8-max; qwen3.7-max is now a legacy model. The moving alias and dated Qwen3.7 snapshots described in our July 28 record remain below only to identify the exact tested route. Do not transfer predecessor results, modality details, or structured-output behavior to Qwen3.8 Max without a fresh matched test.

Qwen3.6 open weights
The current open-weight shortlist includes Qwen3.6-27B and Qwen3.6-35B-A3B. The 27B model lists a vision encoder, 262,144 native context, extension up to 1,010,000 tokens, and Apache 2.0. The MoE model lists 35B total parameters with 3B active and is also released under Apache 2.0.
Older Qwen3 sizes remain useful across hardware tiers, but a current comparison should not omit the Qwen3.6 models designed for agentic coding, multimodal work, and local deployment.
Original Live Evidence: DeepSeek and Qwen
Tested July 28, 2026. We used one synthetic English fixture across both providers. It set a $12,000 budget, a May 15 launch deadline, two vendor delivery dates, and a two-business-day security review. The expected decision was Vendor B: its May 12 delivery leaves two business days before launch, while Vendor A arrives after the deadline. The response had to be valid JSON with one decision, exactly two risks, and one next step.
This is a narrow test of constraint handling, date reasoning, JSON compliance, mode selection, and output budgeting. It is not a coding benchmark or a general quality and speed ranking. Elapsed time includes network and service time; each API row is one run.
| DeepSeek run | Observed result | Tokens and elapsed time |
|---|---|---|
| Chat Instant | Valid JSON, but the vendor decision contradicted the dates | Chat UI; token data unavailable |
| Chat Expert | Correct decision and requested JSON | Chat UI; token data unavailable |
deepseek-v4-flash | Correct | 162 prompt, 205 completion, 127 reasoning; 2,160 ms |
deepseek-v4-pro, 500 max tokens | No final answer; all completion tokens were used for reasoning | 500-token output cap exhausted |
deepseek-v4-pro, 1,600 max tokens | Correct | 162 prompt, 701 completion, 574 reasoning; 12,259 ms |
| Qwen API run | Decision / JSON | Provider-reported tokens | Elapsed | Estimated Singapore list cost |
|---|---|---|---|---|
qwen3.7-flash, default thinking | Correct / valid | 175 prompt, 2,238 completion, 2,103 reasoning | 19,883 ms | $0.0002962 |
qwen3.7-plus, default thinking | Correct / valid | 175 prompt, 2,214 completion, 2,030 reasoning | 41,990 ms | $0.0036124 |
qwen3.7-max, thinking off | Incorrect / valid | 177 prompt, 142 completion, 0 reasoning | 5,711 ms | $0.0015075 |
qwen3.7-max, thinking on | Correct / valid | 175 prompt, 2,353 completion, 2,195 reasoning | 46,814 ms | $0.0180850 |
qwen3.6-flash, thinking on | Correct / valid | 175 prompt, 1,849 completion, 1,731 reasoning | 15,171 ms | $0.0028173 |
Qwen Studio’s Qwen3.7-Plus with Thinking set to Auto also returned the correct decision and valid JSON. The strongest finding is not that one provider is universally better: on this fixture, Qwen3.7 Max was quicker and cheaper with thinking disabled but wrong; enabling thinking made it correct while substantially increasing output tokens and elapsed time. DeepSeek showed the same class of operational risk when Instant was wrong and V4-Pro produced no final answer under a 500-token cap.
The Qwen account’s free quota covered these calls, but the cost column deliberately uses official Singapore list prices so the figures remain comparable. Reasoning tokens are part of completion/output billing. Promotions and free allocations were excluded.
Using DeepSeek’s cache-miss input rate, the successful V4-Flash run cost approximately $0.0000801 and the successful 1,600-token V4-Pro run approximately $0.0006803. Qwen3.7 Flash’s correct run was about $0.0002962. This is why cost per successful task is more informative than list price alone: Qwen3.7 Flash has lower uncached short-prompt rates, but it generated roughly ten times as many completion tokens on this fixture.



Model-list and alias observation
A live GET /models request returned deepseek-v4-flash and deepseek-v4-pro. Although DeepSeek had announced that deepseek-chat and deepseek-reasoner would be retired after July 24, both aliases returned HTTP 200 in our July 28 account test and identified the returned model as deepseek-v4-flash. The reasoner alias returned reasoning content; the chat alias did not.
The authenticated Singapore Qwen endpoint returned HTTP 200 with 151 model IDs. Relevant IDs included qwen3.7-flash, qwen3.7-plus, qwen3.7-max, dated Qwen3.7 snapshots, and Qwen3.6 predecessors. The live console independently exposed Qwen3.7 Flash and the Max May 20 and June 8 snapshots.
Boundary: that is a July 28, 2026 Singapore account observation. It does not override Alibaba’s current public region/catalog documentation.
Do not rely on moving-alias compatibility. Use explicit IDs, log the returned model, and retest after provider changes. Alibaba’s model lifecycle policy gives snapshots shorter notice than mainline aliases, so pin a dated snapshot for reproducibility and monitor retirement notices.
DeepSeek vs Qwen Pricing
There is no single cheapest provider. In the July 28 Singapore rate snapshot, Qwen3.7 Flash had lower uncached input and output list rates than DeepSeek V4-Flash, while DeepSeek had the lower cached-input rate. Actual task cost depends on the current model, region, output length, thinking tokens, cache hits, retries, and correctness. On our dated fixture, DeepSeek V4-Flash was cheaper because its correct answer used far fewer completion tokens.
| Model and effective scope | Input per 1M | Output per 1M | Boundary |
|---|---|---|---|
| DeepSeek Flash — through Aug 16, 15:59 UTC | $0.0028 cached / $0.14 uncached | $0.28 | Current rate |
| DeepSeek Pro — through Aug 16, 15:59 UTC | $0.003625 cached / $0.435 uncached | $0.87 | Current rate |
| DeepSeek Flash — off-peak after cutover | $0.007 cached / $0.22 uncached | $0.66 | All non-peak UTC times |
| DeepSeek Flash — peak after cutover | $0.014 cached / $0.44 uncached | $1.32 | 01:00–04:00 and 06:00–10:00 UTC |
| DeepSeek Pro — off-peak after cutover | $0.022 cached / $0.66 uncached | $1.98 | All non-peak UTC times |
| DeepSeek Pro — peak after cutover | $0.044 cached / $1.32 uncached | $3.96 | 01:00–04:00 and 06:00–10:00 UTC |
| Qwen3.7 Flash — July 28 Singapore snapshot | $0.03 | $0.13 | Up to 32K |
| Qwen3.7 Flash — July 28 Singapore snapshot | $0.10 | $0.40 | Above 32K to 256K |
| Qwen3.7 Flash — July 28 Singapore snapshot | $0.20 | $0.80 | Above 256K to 1M |
| Qwen3.7 Plus — July 28 Singapore snapshot | $0.40 | $1.60 | Up to 256K |
| Qwen3.7 Plus — July 28 Singapore snapshot | $1.20 | $4.80 | Above 256K to 1M |
| Qwen3.7 Max — July 28 Singapore snapshot | $2.50 | $7.50 | Up to 1M |
Current Qwen3.8 Max list pricing checked August 16 is CNY 12 input and CNY 36 output per 1M tokens in China (Beijing) and for Global deployments in US (Virginia), Germany (Frankfurt), and Japan (Tokyo). Singapore’s International deployment is CNY 14.988 input and CNY 44.965 output. Alibaba’s catalog lists Qwen3.8 Max in Beijing, Singapore, Tokyo, Frankfurt, and US (Virginia); verify the exact region and deployment scope before budgeting. See Alibaba Cloud’s current pricing table and DeepSeek’s current pricing page. Promotions and free quotas remain excluded. The July 28 measured costs above remain calculated from the rates in force on that test date.
Coding, Reasoning, Tools, and Long Context
Coding and reasoning
DeepSeek V4-Pro and Qwen3.8 Max are the current hosted higher-tier candidates. The July 28 Qwen3.7 Max and Flash runs on this page remain predecessor evidence, not current-model results. For local coding, Qwen3.6-27B and 35B-A3B remain more practical comparison points than the much larger DeepSeek V4 weights.
Score repository tasks with compilation and tests. Separate code generation, debugging, refactoring, terminal use, and long-horizon agents. A single LiveCodeBench or terminal score cannot predict all five.
Structured output and tools
Both platforms document function or tool calling. On DeepSeek, Chat Completions uses json_object and caller-executed functions, while the stateless Responses API supports both V4 models, json_schema, and server-side web_search. Current Qwen3.8 Max documentation lists thinking, function calling, built-in tools, and structured output. The Qwen3.7 Plus and Flash behavior described by the July 28 test remains predecessor evidence. Test valid calls, missing required arguments, invalid enum values, parallel calls, retries, and multi-turn state on the exact current endpoint.
Long context
DeepSeek V4 and selected hosted Qwen models list 1M context. Qwen3.6-27B lists 262,144 native context with an extension path above 1M. These are not equivalent claims. Test native and extended modes separately, and place answerable evidence at different positions inside 32K, 128K, and 256K contexts before scaling further.
Multimodal Work: Qwen Has the Documented Advantage
July 28 predecessor boundary: Alibaba’s vision documentation then listed Qwen3.7 Plus and Qwen3.7 Flash for text, images, and video, with 1M context and up to 64K output. The dated qwen3.7-max-2026-06-08 snapshot was multimodal, while the moving Max alias pointed to the May 20 text-only snapshot. Those Qwen3.7 Max details are now legacy and should not be transferred to Qwen3.8 Max. DeepSeek’s current official V4 API comparison centers on text generation, reasoning, JSON, and tools rather than native image or video input.
Use Qwen first for screenshot analysis, chart extraction, visual document workflows, and video understanding. This is a documented feature-fit conclusion, not a claim that Qwen won an image-quality benchmark we did not run. Each modality still needs its own accuracy, latency, token, and failure tests.
Local Deployment and Licensing
Qwen’s current open-weight models cover more practical hardware tiers. Qwen3.6-27B is dense; Qwen3.6-35B-A3B activates 3B of its 35B total parameters. Both are dramatically smaller in total parameters than DeepSeek V4-Flash at 284B and V4-Pro at 1.6T.
DeepSeek V4 lists MIT licensing, while the Qwen3.6 open models list Apache 2.0. Both are familiar permissive licenses, but verify the exact repository and included components. A model repository can depend on code, tokenizer assets, datasets, or third-party tools with additional terms.
Before calling any model “local,” record the quantization, disk size, RAM or VRAM, runtime, context, tokens per second, cold start, and output quality loss. Use the local DeepSeek guide and GGUF vs Safetensors guide for implementation details instead of duplicating them here.
Privacy and Data Handling
Privacy depends on the product layer. Alibaba Cloud’s Model Studio privacy notice treats prompts and outputs as customer content and says Alibaba will not use them to develop or improve Model Studio models without separate consent. It also describes encryption and security controls, but this is not a zero-retention promise.
Singapore region selection does not mean every inference operation stays physically in Singapore. Alibaba says stored request data is in Singapore while inference may use international nodes outside mainland China; transient inference-node data is not persisted. The consumer Qwen Studio service has separate privacy terms that allow some de-identified content and feedback to improve services and models. Do not transfer the API policy to the consumer chat product.
DeepSeek’s hosted service follows DeepSeek’s policy. Self-hosted DeepSeek or Qwen follows the storage, logging, operators, and network controls of the environment you manage. Compare the exact provider, region, plan, retention, abuse-monitoring logs, cross-border processing, and contractual terms before processing confidential data.
Reproducible Account-Based Test Plan
The live fixture above is reproducible, but one task is only a smoke test. A reliable production decision needs a larger matched suite, repeated runs, and the same region and scoring rules. Self-hosted Qwen or DeepSeek results also require suitable hardware. Use this expansion plan:
- Record exact model IDs, snapshots, region, endpoint, mode, prompt format, maximum output, and date.
- Compare V4-Flash with an appropriate current Qwen economy model, then V4-Pro with Qwen3.8 Max for the higher tier. Keep Qwen3.7 and Qwen3.6 results only as dated predecessor baselines.
- Run 10 coding tasks with tests, 10 reasoning/extraction tasks, five JSON schemas, and five tool-call scenarios.
- Test multi-turn reasoning-state preservation and measure the token overhead of passing prior reasoning or summaries.
- Run 32K, 128K, and 256K retrieval fixtures with facts at multiple positions. Extend farther only after the lower tiers pass.
- Run Qwen3.7 Plus image and visual-document tasks separately. Mark unsupported DeepSeek modalities as N/A, not zero.
- Use at least three repetitions for stochastic tasks and report correctness, schema validity, tool errors, retries, token use, latency, and region-specific official cost per success.
- Publish sanitized prompts, expected answers, scoring code, and raw result fields without keys, balances, workspace IDs, account details, or private content.
The DeepSeek evaluation framework can organize the test cases. Related implementation references include the DeepSeek API guide, JSON output guide, tool-call guide, and context-caching guide.
Practical Decision Framework
- High-volume text: start with DeepSeek V4-Flash and calculate cost with real input/output ratios.
- Hard hosted coding or reasoning: benchmark DeepSeek V4-Pro against current Qwen3.8 Max.
- Images, video, or visual documents: start with Qwen3.7 Plus.
- Alibaba Cloud stack: Qwen offers the more direct platform fit, regional endpoints, and built-in services.
- Moderate self-hosting infrastructure: test Qwen3.6-27B or 35B-A3B before attempting the much larger DeepSeek V4 weights.
- Mixed workload: route low-cost text to DeepSeek, multimodal requests to Qwen, and keep both as tested fallbacks.
Use the DeepSeek model catalog, V4 guide, and pricing page for deeper DeepSeek-specific details. This comparison should remain the decision page rather than repeating those complete guides.
Update Log
- July 28, 2026: rebuilt the comparison around DeepSeek V4 and the then-current Qwen3.7 Max, Plus, and Flash lineup; corrected Singapore pricing, model snapshots, multimodal support, structured-output caveats, privacy, endpoint, and lifecycle details.
- August 14, 2026: added the public Qwen region/catalog boundary and updated only the current DeepSeek layer for Pro-0813 GA, Flash-0731 public beta, Responses on both models, endpoint-specific structured output/search, reasoning mapping, and the August 16 price cutover. Dated account tests and costs remain unchanged.
- Original testing: ran one matched synthetic fixture in DeepSeek Chat, the DeepSeek API, Qwen Studio, and five Qwen API configurations; added sanitized screenshots, tokens, elapsed time, correctness, and list-cost estimates.
- Limitations: API timings are single runs, local Qwen3.6 and DeepSeek V4 weights were not benchmarked, and the fixture does not substitute for a coding, multimodal, or long-context suite.
FAQ: DeepSeek vs Qwen
Is DeepSeek better than Qwen?
DeepSeek is the stronger first choice for low-cost hosted text workloads in the listed price examples. Qwen is the stronger platform fit for multimodal input, Alibaba Cloud integration, and smaller current open-weight models. Quality needs a matched benchmark.
Which is better for coding?
Test DeepSeek V4-Pro and current Qwen3.8 Max for hosted coding agents. The Qwen3.7 Max result on this page is a July 28 predecessor test, not a Qwen3.8 result. For a local coding model, Qwen3.6-27B or 35B-A3B is more practical than DeepSeek V4 for many teams. Use compilation, unit tests, tool errors, and accepted patches rather than one benchmark score.
Which is cheaper?
In the July 28 Singapore snapshot, Qwen3.7 Flash had lower uncached short-prompt input and output list rates, while DeepSeek V4-Flash had a lower cache-hit input rate. DeepSeek V4-Flash cost less on that one fixture because it generated far fewer output tokens. For a new budget, use the current Qwen3.8 Max regional prices above and compare cost per correct result.
Which is better for multimodal work?
Qwen has the documented feature advantage. Qwen3.7 Plus supports text, image, and video input with a 1M context window. DeepSeek’s current official V4 API comparison is text-focused.
Which is easier to run locally?
Qwen is usually the more practical current starting point because Qwen3.6 includes 27B and 35B-A3B open weights. DeepSeek V4-Flash and Pro are far larger in total parameters. Actual feasibility still depends on quantization, runtime, context, and hardware.
Do DeepSeek and Qwen both support 1M context?
DeepSeek lists 1M for both V4 API models, and Alibaba lists 1M for current Qwen3.8 Max. The July 28 snapshot also documented 1M for selected Qwen3.7 hosted models. Qwen3.6-27B lists a smaller native context with an extension path above 1M. Test retrieval rather than comparing maximum numbers alone.
Are both model families open source?
Open-weight is the more precise general term. DeepSeek V4 repositories list MIT, while current Qwen3.6 open-weight repositories list Apache 2.0. Hosted Qwen3.7 aliases should not be described as downloadable merely because other Qwen models publish weights.
What did the live Qwen test show?
Qwen Studio’s Qwen3.7-Plus and four thinking-enabled API configurations returned the correct decision and valid JSON. Qwen3.7 Max with thinking disabled returned valid JSON but the wrong decision; enabling thinking corrected it with much higher token use and elapsed time. These are single-fixture observations, not general rankings.
Official Sources
- DeepSeek Models and Pricing
- DeepSeek API Updates
- DeepSeek Responses API
- DeepSeek V4 Preview Release
- Alibaba Cloud Model Studio Public Model Catalog
- Newly Released Qwen Models
- Current Alibaba Cloud Model Studio Pricing
- Qwen Text-Generation Capabilities
- Qwen Visual Understanding
- Model Studio Regions, Deployment Scopes, and Endpoints
- Model Studio Lifecycle and Retirement Policy
- Model Studio Security and Privacy Notice
- Qwen Studio Privacy Policy
- Official Qwen3.6 Repository
- Qwen3.6-27B Model Card
- Qwen3.6-35B-A3B Model Card
