DeepSeek Tests: What We Ran, When, and What We Found

This page collects every hands-on test we have run against DeepSeek, with the date, the mode, the rubric score and a link to the full write-up. Nothing here is a vendor benchmark or a re-published number. Each score comes from a run we recorded ourselves, and each linked page reports what failed as well as what worked.

The July 2026 task batch

Nine task tests share one protocol, recorded on 29 July 2026:

  • Product: DeepSeek Chat, Instant mode.
  • Inputs: fixed synthetic, English-only datasets and briefs, written for the test so no private or client data was exposed.
  • No web search enabled, so the model answered from the supplied context only.
  • A fixed rubric per task, scored item by item, with a single recorded run of the prompt published on the page.
TestScorePercentFull write-up
How to Use DeepSeek AI for Translation15/15100%Read the test
DeepSeek for Excel Analysis13/13100%Read the test
DeepSeek for Obsidian and Personal Knowledge Management24/2596.0%Read the test
DeepSeek for Market Research15/1693.8%Read the test
DeepSeek for Data Analysis25/2792.6%Read the test
DeepSeek for Productivity17/2085.0%Read the test
DeepSeek for Game Developers and Indie Studios24/2982.8%Read the test
DeepSeek for Entrepreneurs, Freelancers and Small Business Owners13/1681.3%Read the test
DeepSeek for Students13/1776.5%Read the test

Scores range from 76.5% to 100%. The low end is not hidden: the student test lost points on citation accuracy, and the game-development test lost points because the model added setting details and production estimates that the brief never supplied. Those failures are written up on their own pages.

Comparison and endpoint tests

Two tests use different protocols and are listed separately so the numbers are not mixed with the task batch:

TestDateMethodFull write-up
DeepSeek vs ChatGPT7 August 2026 snapshotUnified comparison protocol v1.0. API lane: deepseek-v4-flash High versus gpt-5.6-terra High. UI and API attempts kept in separate tables with separate denominators. Sanitized observations, isolated API aggregates and blinded citation adjudications retained locally with SHA-256 hashes.Read the test
DeepSeek API balance endpoint25 July 2026Live test against the official endpoint using three privacy-safe requests.Read the test

How to read these scores

A rubric score answers a narrow question: on this task, with these inputs, how much of what we asked for did the model deliver? It is not a general capability rating, and it does not transfer between tasks. Four limits are worth stating plainly:

  • Synthetic inputs. The datasets and briefs were written for the test. Real, messy production data is harder.
  • One recorded run. Each published result is a single run of a fixed prompt, not an average over repeated runs.
  • One mode. The July batch used Instant mode only. Expert mode and the reasoning levels were not part of it.
  • English only. No conclusion here should be read as applying to other languages.

Why we publish this

Model behaviour changes with every release, and a claim about what DeepSeek can do is worth nothing without a date attached to it. Keeping the runs, the dates and the failures in one place is what lets a reader check us rather than trust us. When a test is re-run against a newer model, the page is updated and the date on it changes.

Privacy and cookie settings