Can DeepSeek cite sources accurately? Across 48 English questions in 30 canonical scored chats, our answer was not reliable enough for unverified use. The research combines a July 2026 global-domain baseline with a separately scored August 2026 country expansion covering official sources from the United States and India. Two additional incomplete capture attempts were excluded under the documented rerun rule.
In the July baseline, 76 of 96 citation occurrences resolved to valid substantive sources or were confirmed manually, for 79.1667% URL validity. The stricter claim-level review found 118 supported-edge points across 169 citation-claim edges, or 69.8225% support precision. Weighted citation completeness was 75/132, or 56.8182%. DeepSeek passed 4 of 24 questions and 0 of six batches, producing a final macro score of 61.3899/100.
Those figures measure different things. A valid URL can lead to the wrong evidence. A relevant page can fail to support the sentence beside it. A citation can support one clause while leaving another ungrounded. An otherwise correct answer can also omit a material date, denominator, exception, version, or scope qualifier.
Benchmark scope: We tested the official DeepSeek consumer web chat with Search on and DeepThink off. The July baseline used six fresh chats, each containing one exact four-question batch. The August country expansion used 24 additional fresh chats, one exact question per chat. Only the first completed answer was scored, except for two documented operator-side capture failures that were excluded and rerun once under the same prompt. The results do not establish the behavior of the DeepSeek API, open-weight models, third-party hosts, no-search chat, other accounts, other regions, or future product versions.
For broader product context, start with our independent DeepSeek guide. This article stays deliberately narrow: it measures whether citations returned by two dated DeepSeek Search protocols were reachable, correctly placed, authoritative, complete, and supportive. The complete question packs, results, methodology, and machine-readable files are available in the public DeepSeek citation-audit repository.
DeepSeek citation accuracy: the short answer
The July baseline, deepseek-web-citation-audit-v1 version 1.0.0, used 24 English questions covering health guidance; Earth, space, and climate science; public data and economics; law, regulation, and security policy; web and software standards; and history and world heritage.
The overall result was FAIL under the benchmark’s predeclared hard gates.
| Measure | Audited result | Denominator | What it means |
|---|---|---|---|
| Overall macro score | 61.3899/100 | 24 question scores | Mean of the final question-level numeric scores |
| Questions passed | 4/24 | 24 questions | Questions meeting the numeric threshold and every question hard gate |
| Batches passed | 0/6 | Six four-question batches | Batches meeting the numeric threshold and every batch hard gate |
| Weighted factual accuracy | 115.5/132 (87.5%) | 132 gold-claim weight points | Whether required facts, numbers, dates, scope, and qualifiers were correct |
| Citation URL validity | 76/96 (79.1667%) | 96 citation occurrences | Whether each displayed citation reached the intended substantive source |
| Citation support precision | 118/169 (69.8225%) | 169 citation-claim edges | Whether cited sources entailed the atomic claims a reader would associate with them |
| Weighted citation completeness | 75/132 (56.8182%) | 132 gold-claim weight points | Whether correct required claims had adjacent, valid, fully supporting citations |
| Weighted authority coverage | 67.5/132 (51.1364%) | 132 gold-claim weight points | Whether full support came from sources appropriate to the claim |
| Citation placement precision | 70/96 (72.9167%) | 96 citation occurrences | Whether citations were adjacent to at least one supported claim |
| Preferred-source hit rate | 12/24 (50%) | 24 questions | Questions using a frozen preferred source or an exact official equivalent |
Most important strength: Weighted factual accuracy reached 87.5%, and Q10, Q11, Q13, and Q17 combined accurate answers with complete and authoritative support.
Most consequential failure: Only 56.8182% of required claim weight received full citation-completeness credit. Four answers had no rendered citations, and two more had citations but no URL that could earn validity credit.
Practical recommendation: Use DeepSeek Search to discover possible sources, not as a final authority. Open the exact URL and verify every material clause before relying on, publishing, or citing the answer.
August 2026 U.S. and India country expansion
On August 6, 2026, we added 24 new questions: 12 tied to verified U.S. official sources and 12 tied to verified Indian official sources. Each canonical result used its own fresh chat under the same visible consumer settings: Instant mode, Search on, DeepThink off, English only, and the first completed answer scored. Two incomplete operator-side captures were excluded and rerun once as documented below.
We report this expansion separately from the July baseline. The baseline submitted four questions per chat and used a full batch gate; the expansion submitted one question per chat and has no batch score. Combining their macro scores would hide that protocol difference.
| Country subset | Questions passed | Macro score | Factual accuracy | URL validity | Support precision | Completeness | Authority coverage | Placement |
|---|---|---|---|---|---|---|---|---|
| United States | 4/12 | 74.8202/100 | 83/93 (89.2473%) | 103/115 (89.5652%) | 147.5/180 (81.9444%) | 70/93 (75.2688%) | 61/93 (65.5914%) | 102/115 (88.6957%) |
| India | 0/12 | 42.1966/100 | 68/99 (68.6869%) | 61/80 (76.25%) | 91.5/129 (70.9302%) | 34/99 (34.3434%) | 25/99 (25.2525%) | 60/80 (75%) |
No distinct country-subset pass rule was frozen before the live run. The canonical outcomes are the question-level pass counts and the metrics above. The machine-readable score files also provide a descriptive FAIL label by applying the inherited V1 hard-gate logic, but it should not be read as a separately preregistered country-aggregate outcome.
The U.S. subset passed four questions: constitutional qualifications, post-1977 copyright duration, CISA Secure by Design principles, and USGS earthquake terminology. The subset missed the inherited aggregate thresholds. One DOT answer omitted material refund conditions; the EEOC answer replaced the special age-discrimination filing condition; and the FEMA answer misstated or overgeneralized waiting-period exceptions. Several other answers were factually strong but failed support, authority, or unresolvable-link gates.
The India subset passed no question. The most serious error appeared in the UIDAI e-KYC answer: it said biometric authentication was prohibited, while current UIDAI guidance says e-KYC authentication is performed using OTP and/or biometric authentication. Other material failures included reversing the cancellation-charge condition in the E-Commerce Rules, reporting non-final Census 2011 population figures instead of the frozen final values, and omitting required probabilities and error bounds from the dated April 13, 2026 IMD monsoon forecast.
These country labels describe the official-source subject matter, not the physical location of the user, browser, retrieval service, or model backend. The consumer interface did not expose those systems, so this is not a latency or geographic-routing benchmark.

What changed in version 2.0.0
- One question was submitted per fresh chat to isolate context and make question-level reruns auditable.
- The U.S. pack froze 12 questions, 93 weighted claim points, and official sources from bodies including Congress.gov, the Copyright Office, EEOC, DOT, NIST, CISA, FDA, USDA, TreasuryDirect, USGS, and FEMA.
- The India pack froze 12 questions, 99 weighted claim points, and official sources including India Code, Consumer Affairs, DICGC, UIDAI, ECI, ISRO, Census of India, IMD, BIS, FSSAI, and MeitY.
- The same dimensions remained separate: factual accuracy, citation URL validity, claim-level support, completeness, authority, placement, and format.
- The inherited format rubric awarded one point for an exact question-ID heading even though the V2 prompts did not request one. Both subsets lost that one component consistently, but no separate severity deduction was added for an instruction the prompts never gave. The mismatch did not cause any otherwise passing question to fail.
Download the data and reproduce the audit
The public GitHub repository contains the frozen question packs, gold claims, methodology, scored results, data dictionaries, checksums, and a CC BY 4.0 license for the original release data and documentation.
- Download the July question results as CSV
- Download the July full dataset as JSON
- Download the U.S. and India question results as CSV
- Download the country expansion results as JSON
- Review the frozen U.S. and India benchmark packs
Private chat URLs, account identifiers, and raw interface snapshots are intentionally excluded. The public files preserve prompts, gold claims, official source records, scores, exact denominators, findings, input hashes, and protocol notes without exposing the signed-in session.
What counts as an accurate DeepSeek citation?
Citation accuracy is not one binary property. We used seven checks because each catches a different failure.
| Dimension | Pass condition | Why it matters |
|---|---|---|
| URL validity | The citation reaches the intended substantive source within five redirects, or a normal browser confirms a blocked official source | A broken, deceptive, or unresolvable link cannot be verified |
| Target identity | The final page is the source described by the citation label | A real URL can still point to the wrong article, document, or publisher |
| Claim support | The source directly entails the adjacent atomic claim | A topically related page may not prove the model’s statement |
| Completeness | Every correct required claim has an adjacent valid citation with full support | A few good citations do not ground an entire answer |
| Source authority | The source is appropriate for the exact claim | An original law, standard, dataset, or agency record is stronger than a derivative summary within its remit |
| Freshness and version | The source matches the date, revision, or historical frame requested | A formerly correct page can be wrong for a current-rule question |
| Placement | A reasonable reader can identify which claim the citation supports | A detached or decorative citation makes claim-level verification difficult |
The distinction between answer correctness and citation quality follows the core logic of the ALCE citation-evaluation benchmark. We also split compound sentences into independently checkable claims, reflecting the fine-grained citation problem examined by ALiiCE. Our practical method is documented under our editorial and source standards. It is an independent adaptation for live consumer web citations, not an official ALCE or ALiiCE replication.
A valid page is not automatically valid evidence
Suppose an answer says:
Agency X adopted Rule Y in 2025, and the rule requires Z.
One source might confirm the agency and adoption date but describe a different requirement. The URL can return a successful page, the publisher can be authoritative, and the citation can still fail the requirement claim.
The same boundary applies to HTTP status checks. Under RFC 9110, 200 OK indicates that the request succeeded. It does not establish that the returned page is the intended document, that the information is current, or that the page supports DeepSeek’s adjacent statement.
How the final edge denominator was established
The first scoring pass created one nearest-claim support judgment per citation occurrence. That was useful as a conservative checkpoint, but 96 occurrences could not serve as the final claim-edge denominator because the frozen rubric requires an edge for every atomic claim that a reasonable reader would associate with a citation.
The reconciliation expanded compound citation clusters to 169 citation-claim edges. It included gold-aligned atomic response claims plus two unambiguous extra factual claims, and it split Q06’s river and fresh-lake percentages because either number could be true without the other. B01-B03 received source-specific edge labels; B04-B06 used the expanded atomic edge map while preserving the primary support label at each citation occurrence. The resulting 118/169 figure is the canonical conservative reconciliation for this publication. A de novo second page-level review of B04-B06 could change individual labels, but it could not justify returning to a one-edge-per-occurrence denominator.
The source-specific review changed four earlier inherited-label decisions. It treated the Q05 Horizons launch-date edge as partial and its phase-name edge as unsupported; found that the Q06 USGS table contradicted the river percentage while supporting the fresh-lake percentage; credited the Q09 Census source for both resident-population exclusion and home-state apportionment treatment; and denied Q12 support for an active-duty permanent-change-of-station scope that its cited rate document did not establish.
How we designed the July 24-question citation benchmark
The questions, prompts, gold claims, source preferences, scoring formulas, thresholds, and hard gates were frozen before the first live submission. Changing them after seeing an output would create a new benchmark version.
The frozen design contained:
- 24 scored questions;
- six topic batches;
- four questions per batch;
- 101 required atomic claims;
- 132 total required-claim weight points;
- 28 preferred Tier A source records;
- one exact prompt per batch;
- one fresh chat per batch;
- Search on and DeepThink off;
- first completed responses only;
- no regeneration, repair prompt, clarification, or substantive follow-up.
The 28 preferred records were original agencies, standards bodies, laws, official datasets, mission records, treaty text, or institutional primary sources. An equivalent official source could receive the same Tier A rating even when its URL differed from the frozen preferred page.

Six batches, six evidence domains
| Batch | Questions | Domain | Examples of required evidence |
|---|---|---|---|
| B01 | Q01-Q04 | Health guidance | CDC activity and sleep guidance, FDA caffeine guidance, NIH vitamin D guidance |
| B02 | Q05-Q08 | Earth, space, and climate science | NASA mission records, USGS water estimates, ozone authorities, IPCC findings |
| B03 | Q09-Q12 | Public data and economics | Census definitions, BLS CPI scope, Federal Reserve framework, historical IRS rates |
| B04 | Q13-Q16 | Law, regulation, and security policy | Fair-use factors, CAN-SPAM requirements, final NIST password rules, GDPR Article 17 |
| B05 | Q17-Q20 | Web and software standards | WCAG contrast, HTTP 429, robots exclusion, official Python semantics |
| B06 | Q21-Q24 | History and world heritage | Library of Congress, National Archives, UNESCO, and National Park Service records |
The questions were designed to expose recurring citation problems: numbers without units or denominators, obsolete revisions, partly supported compound sentences, legal rights without exceptions, statistics without population definitions, and technically correct claims sourced to weak derivative pages.
Exact response contract
Every batch prompt required DeepSeek to:
- answer all four questions;
- use the exact question ID as an H2 heading;
- stay within each question’s sentence range;
- cite every factual sentence inline with clickable citations;
- prefer original agencies, official standards, treaty text, mission records, and institutional primary sources;
- avoid search-results pages, AI-generated pages, citation aggregators, and unsourced snippets;
- add no preface, bibliography, conclusion, or advice outside the four answers.
The experiment’s unit was one first completed response to one four-question batch prompt. Every batch started in a fresh chat, so previous answers could not become context for later topics.
Controlled interface conditions
| Setting | Recorded value |
|---|---|
| Product surface | Official DeepSeek consumer web chat |
| Test date | July 26, 2026 |
| Test window | 11:27:38 a.m. to 11:41:21 a.m. PDT (18:27:38 to 18:41:21 UTC) |
| Visible mode label | Instant |
| Search | On for all six batches |
| DeepThink | Off for all six batches |
| Language | English only |
| Account scope | One signed-in consumer web session; account plan and identifier were not exposed or published |
| Chats | Six fresh chats |
| Regenerations | 0 |
| Corrective follow-ups | 0 |
| Recorded protocol deviations | 1 |
The consumer interface did not expose the underlying model identifier or version, backend routing, retrieval index or provider configuration, account plan, region, browser, or interface locale. We do not infer those values. A browser test should not be presented as a benchmark of a named API model unless the interface establishes that identity.
DeepSeek’s official app announcement lists Web Search among the product’s features. That documentation does not guarantee identical retrieval behavior across every interface, account, region, or date.
Consumer Search and API retrieval are separate product questions. Our DeepSeek API Web Search guide documents that boundary; this audit does not infer API capabilities from buttons visible in the chat interface.
How every answer and citation was scored

Scoring began after all six first responses and their citation inventories were frozen. Each answer was split into atomic claims before citation quality was judged.
| Section | Maximum points | What it measured |
|---|---|---|
| Factual accuracy | 35 | Required claims against the frozen gold answer |
| URL validity | 15 | Citation occurrences reaching intended substantive sources |
| Claim support | 20 | Source entailment across citation-claim edges |
| Citation completeness | 15 | Correct required claims with full adjacent support |
| Source authority | 10 | Best fully supporting source tier for each required claim |
| Format and placement | 5 | Exact heading, length, prohibited extras, and placement precision |
A question needed at least 85/100, at least 90% factual accuracy, 100% URL validity, at least 90% support precision, 100% completeness, at least 80% authority coverage, no fabricated or unresolvable citation, no critical event, and no material contradiction of a required claim.
A batch needed at least 90/100, all four questions to pass, no invalid or unverifiable citation occurrence, no critical event, and all exact question headings. The full benchmark needed at least 90/100, all six batches to pass, at least 95% overall factual accuracy, 100% URL validity, at least 95% support precision, at least 98% completeness, at least 90% authority coverage, no fabricated or unresolvable citation, no critical event, and no contamination between fresh chats.
The numeric score did not override hard gates. That is why Q03 scored 96.3333 but failed its 90% support gate, and why Q13 scored 97 yet still passed despite a small format deduction.
Teams adapting this method to their own applications can use the DeepSeek evaluation framework to separate datasets, evaluators, thresholds, and regression runs. The benchmark files in this research release remain the source of truth for the numbers reported here.
URL states were kept separate
- Valid: the citation resolved to the intended substantive page or official document.
- Valid manual: automated access was blocked, but independent retrieval confirmed the intended publisher, page, and relevant content.
- Invalid: the link was malformed, broken, deceptive, unrelated, a soft 404, a search-results page, or a source-free internal citation page.
- Unverifiable: neither automated nor manual inspection could establish the target and substantive content.
Redirect chains were followed for no more than five redirects. Raw links were preserved, while normalized links were used only for duplicate detection. Automated inaccessibility alone was not labeled fabrication.
Source authority was scored only after support
| Tier | Source type | Score |
|---|---|---|
| A | Original law, regulation, standard, official dataset, mission record, or authoritative agency page | 1.00 |
| B | High-quality institutional synthesis or scholarly secondary source | 0.75 |
| C | General secondary source, commercial explainer, news report, or encyclopedia | 0.25 |
| D | User-generated, AI-generated, scraped, low-accountability, citation-aggregation, or search-results source | 0.00 |
A prestigious domain did not earn automatic authority credit. The page first had to support the claim.
Results across all six citation batches

| Batch | Domain | Macro score | Questions passed | Factual accuracy | URL validity | Support precision | Placement | Completeness | Authority coverage | Result |
|---|---|---|---|---|---|---|---|---|---|---|
| B01 | Health guidance | 67.5040 | 0/4 | 83.3333% | 88.8889% | 16.5/26 (63.4615%) | 15/18 (83.3333%) | 72.2222% | 52.7778% | FAIL |
| B02 | Earth, space, and climate science | 43 | 0/4 | 69.0476% | 75% | 14.5/25 (58%) | 11/16 (68.75%) | 42.8571% | 42.8571% | FAIL |
| B03 | Public data and economics | 82.4583 | 2/4 | 91.3043% | 85.7143% | 18/23 (78.2609%) | 11/14 (78.5714%) | 82.6087% | 82.6087% | FAIL |
| B04 | Law, regulation, and security policy | 76.9048 | 1/4 | 94.4444% | 100% | 45/51 (88.2353%) | 20/21 (95.2381%) | 81.4815% | 73.1481% | FAIL |
| B05 | Web and software standards | 60 | 1/4 | 94.7368% | 47.619% | 20/35 (57.1429%) | 10/21 (47.619%) | 31.5789% | 22.3684% | FAIL |
| B06 | History and world heritage | 38.4722 | 0/4 | 89.5833% | 83.3333% | 4/9 (44.4444%) | 3/6 (50%) | 25% | 25% | FAIL |
B04 is the clearest warning against using link reachability as a proxy for answer quality: all 21 citation occurrences were URL-valid, yet the batch still failed with one passing question, 81.4815% completeness, and 73.1481% authority coverage. B05 shows the opposite pattern in the same run: factual accuracy was 94.7368%, but only 10 of 21 citation occurrences were valid and the batch’s completeness was 31.5789%.

All 24 question results
The table reports final reconciled scores and hard-gate results separately. A high score cannot turn a failed hard gate into a pass.
| Question | Topic | Final score | Result | Most important audited finding |
|---|---|---|---|---|
| Q01 | Adult physical activity | 62 | FAIL | The activity totals and smaller-chunk advice were right, but the all-major-muscle-groups qualifier was missing and one citation was unverifiable. |
| Q02 | Caffeine | 38.6825 | FAIL | The 400 mg framing was preserved, but the coffee-cup conversion and caution population were materially incomplete. |
| Q03 | Vitamin D | 96.3333 | FAIL | All NIH intake values were accurate and complete; one redundant Mayo citation only partially supported the combined age-band sentence, leaving support precision at 80%. |
| Q04 | Adult sleep | 73 | FAIL | The three sleep bands were correct, but one of two citations was unverifiable and the surviving source was a general secondary explainer. |
| Q05 | Voyager 1 | 54 | FAIL | Launch, targets, and August 2012 were right, but the required NASA distinction was replaced with a mission phase name; the Horizons date edge was only partial and its phase-name edge was unsupported. |
| Q06 | Earth’s water | 44 | FAIL | The large reservoirs were accurate; the tiny rivers-and-lakes fraction was materially misstated, the USGS source contradicted the river value, and half the citation occurrences were broken. |
| Q07 | Ozone | 53 | FAIL | Ozone composition and depletion chemistry were good, but the regional recovery timetable did not match the frozen benchmark. |
| Q08 | IPCC warming findings | 21 | FAIL | Most temperature values were right, but the central attribution wording was weakened and the answer contained no live citation. |
| Q09 | 2020 Census population | 78.3333 | FAIL | The resident count, date, scope, and growth were correct, and expanded source support was 100%; the answer’s overseas home-state allocation wording remained factually incomplete. |
| Q10 | Consumer Price Index | 100 | PASS | The CPI basket, substitution distinction, and non-cost-of-living qualifier matched the BLS benchmark with official citations. |
| Q11 | Federal Reserve inflation goal | 100 | PASS | The longer-run 2% PCE framing and measure distinction were accurate, complete, and supported by Federal Reserve sources. |
| Q12 | 2024 IRS mileage rates | 51.5 | FAIL | All three mileage rates were right, but two KPMG links were dead and the cited rate document did not establish the military permanent-change-of-station qualifier. |
| Q13 | Fair use | 97 | PASS | All four statutory factors and the case-by-case rule were accurate and supported; only the requested sentence range was missed. |
| Q14 | CAN-SPAM | 85 | FAIL | CAN-SPAM duties and timing were largely correct, but the advertising-identification exception was omitted. |
| Q15 | NIST password rules | 47.3333 | FAIL | Several password rules were directionally right, but citation support was weak and the named final-revision requirement was not met cleanly. |
| Q16 | GDPR erasure | 78.2857 | FAIL | The grounds and exception categories were mostly accurate, but a legally material necessity limit and the requested sentence range were missed; reconciled support was 88.0952%. |
| Q17 | WCAG contrast | 100 | PASS | Contrast ratios, large-text thresholds, and exemptions were accurate and fully cited; the official W3C link used a WCAG 2.1 understanding path whose relevant ratios match WCAG 2.2. |
| Q18 | HTTP 429 | 38 | FAIL | The semantics were accurate, but every Stack Overflow revision citation remained unverifiable, so none earned support or completeness credit. |
| Q19 | Robots Exclusion Protocol | 38 | FAIL | The robots.txt statements matched RFC 9309, but all six citation occurrences were unverifiable or broken. |
| Q20 | Python lists and tuples | 64 | FAIL | Mutability and singleton syntax were right, but typical access patterns were incomplete and the supporting sources had low authority. |
| Q21 | Jamestown | 84.8889 | FAIL | All required facts were accurate and an official Library of Congress source covered them, but weak placement and one unverifiable news citation reduced precision. |
| Q22 | 19th Amendment | 38 | FAIL | All five required facts matched the benchmark, yet the answer supplied zero citations. |
| Q23 | Great Barrier Reef | 9 | FAIL | Year, criteria, and area were right, but reef and island counts were wrong or over-precise and no citation was supplied. |
| Q24 | Yellowstone | 22 | FAIL | The law, land withdrawal, and 1916 comparison were right; the National Park Service attribution was weakened and the entire answer was uncited. |
The four passing questions were Q10, Q11, Q13, and Q17. In the final reconciliation, Q09’s support gate changed from fail to pass and Q16’s changed from pass to fail; both questions still failed other frozen gates. Q03 illustrates why the gates matter: its factual accuracy, URL validity, completeness, authority coverage, and placement were all 100%, but its expanded claim-edge support precision was 80%, below the required 90%.

Link failures, unsupported claims, and missing citation coverage
These failure classes remain separate. A broken link is not the same as an unsupported claim, and a missing citation is not the same as an incorrect fact.
Invalid or unverifiable citation occurrences
The ten grouped rows below account for all 20 non-valid occurrences: eight invalid and twelve unverifiable. Repeated citation IDs are retained because the benchmark denominator counts every occurrence.
| Citation ID(s) | Question | Displayed label | Public target | Validity | Occurrences | Audit note |
|---|---|---|---|---|---|---|
| B01-CIT05 | Q01 | 5 | Emanate Health GUID page | Unverifiable | 1 | The server returned 403, and independent retrieval did not establish the exact page title or substantive content. |
| B01-CIT17 | Q04 | 4 | DMU Clinic GUID page | Unverifiable | 1 | The server returned 403, and independent retrieval did not establish the exact page title or substantive content. |
| B02-CIT01 | Q05 | 1 | Archived Voyager 1 page | Unverifiable | 1 | A web-application challenge prevented inspection of the archived snapshot. |
| B02-CIT07, B02-CIT09, B02-CIT11 | Q06 | 10 | Water distribution PDF target | Invalid | 3 | The target returned 404 with no substantive document. |
| B03-CIT11, B03-CIT13 | Q12 | 15 | KPMG mileage PDF target | Invalid | 2 | The target returned a 404 page instead of the intended PDF. |
| B05-CIT07, B05-CIT09, B05-CIT11 | Q18 | 2 | Stack Overflow revision-source target A | Unverifiable | 3 | A 403 challenge and cache miss prevented confirmation of the exact revision content. |
| B05-CIT08, B05-CIT10 | Q18 | 6 | Stack Overflow revision-source target B | Unverifiable | 2 | A 403 challenge and cache miss prevented confirmation of the exact revision content. |
| B05-CIT12, B05-CIT14, B05-CIT16 | Q19 | 7 | IETF development-host RFC 9309 target | Unverifiable | 3 | A 403 challenge prevented confirmation of the exact development-host document. |
| B05-CIT13, B05-CIT15, B05-CIT17 | Q19 | 10 | Third-party RFC 9309 PDF target | Invalid | 3 | The target returned 404 with no substantive RFC PDF. |
| B06-CIT03 | Q21 | 13 | Politico Jamestown page | Unverifiable | 1 | A 403 challenge and robots restriction prevented confirmation of the exact page content. |
No fabricated URL was identified. That does not rescue the run: invalid and unverifiable citations still failed the benchmark’s hard gates.
Reachable citations with no support credit
The final edge inventory contained 110 full-support edges, 16 partial-support edges, and 43 no-support edges. The full and partial edges yield 118 support points across 169 edges. Of the 43 no-support edges, 32 came from invalid or unverifiable occurrences listed above. The remaining 11 no-support edges came from nine citation occurrences whose targets were valid or manually confirmed:
| Citation ID(s) | Question | No-support edges | Source | Audited reason |
|---|---|---|---|---|
| B01-CIT11, B01-CIT12 | Q02 | 2 | Mayo Clinic caffeine article | The manually confirmed page did not establish the displayed three-to-four-cup conversion or the separate sensitivity statement. |
| B02-CIT05 | Q05 | 1 | NASA JPL Horizons record | The record described interstellar-mission timing but did not establish the displayed current-phase designation. |
| B02-CIT10 | Q06 | 1 | USGS Earth water distribution | The source contradicted the displayed 0.006% river value; a separate edge from the same occurrence fully supported the fresh-lake value. |
| B02-CIT15 | Q07 | 1 | Cambridge ozone research paper | The paper discussed signal timing but did not establish return to 1980 near-global total-column values under continued policy. |
| B03-CIT14 | Q12 | 1 | Taxpayer Advocate rate document | The document listed the rate but did not establish the active-duty, military-orders, permanent-change-of-station scope. |
| B04-CIT14 | Q15 | 2 | NIST SP 800-63B Rev. 4 Password Updates | The manually confirmed commercial explainer received no support credit for the attached no-composition-rules and blocklist edges. |
| B06-CIT02 | Q21 | 2 | Jamestown classroom materials | At this rendered position, the occurrence received no support credit for the establishment-date and charter edges; separate Library of Congress citations supplied the answer’s full claim coverage. |
| B06-CIT05 | Q21 | 1 | Jamestown classroom materials | At this rendered position, the occurrence received no support credit for the first-permanent-settlement edge; another adjacent Library of Congress citation supplied coverage. |
B01-B03 use source-specific edge evidence. B04-B06 preserve the primary occurrence-level support decisions while expanding edge multiplicity. A de novo second page-level review could revise individual B04-B06 labels, but it would not justify returning to one edge per citation occurrence.
Required claim weight without full completeness credit
Overall, 75 of 132 required-claim weight points received full completeness credit. The remaining 57 points were lost because the answer was wrong or incomplete, no adjacent citation appeared, a citation was invalid or unverifiable, or the adjacent source did not fully support the claim. The table lists every question below 100% completeness.
| Question | Credited weight | Completeness | Main reason for the gap |
|---|---|---|---|
| Q01 | 3/5 | 60% | The strength statement omitted the all-major-muscle-groups qualifier, and one associated citation was unverifiable. |
| Q02 | 2/5 | 40% | The cup conversion was wrong and the caution scope omitted material populations and conditions. |
| Q05 | 3/5 | 60% | The answer omitted the required first-human-made and most-distant-object distinction. |
| Q06 | 3/5 | 60% | The nested rivers-and-lakes fraction was misstated, and three occurrences pointed to a dead PDF. |
| Q07 | 3/6 | 50% | The regional recovery years did not match the frozen benchmark. |
| Q08 | 0/5 | 0% | The attribution language was weakened and the answer rendered no citation. |
| Q09 | 5/7 | 71.4286% | The home-state allocation of eligible overseas counts was incomplete. |
| Q12 | 3/5 | 60% | Two PDF citations were dead and the military moving exception lacked its permanent-change-of-station condition. |
| Q14 | 7/8 | 87.5% | The answer omitted the statutory exception to the advertising-identification duty. |
| Q15 | 3/6 | 50% | The answer softened final-standard wording and relied on weak or obsolete secondary material. |
| Q16 | 6/7 | 85.7143% | The research and archiving exception omitted the impossible-or-seriously-impaired purpose condition. |
| Q18 | 0/5 | 0% | All five citation occurrences were unverifiable. |
| Q19 | 0/6 | 0% | All six citation occurrences were invalid or unverifiable. |
| Q20 | 2/4 | 50% | Required iteration, unpacking, or indexing details were omitted, and source authority was low. |
| Q22 | 0/6 | 0% | The factual answer rendered no citation. |
| Q23 | 0/6 | 0% | Reef and island counts were wrong or over-precise, and the answer rendered no citation. |
| Q24 | 0/6 | 0% | The National Park Service attribution was weakened, and the answer rendered no citation. |
Q03, Q04, Q10, Q11, Q13, Q17, and Q21 received 100% completeness credit. Completeness alone did not make all of them pass: URL validity, support precision, authority, placement, format, and the other hard gates still applied.
Source-tier distribution
The occurrence denominator is 96; the unique-URL denominator is 50. These counts describe the class of cited sources, not whether those sources supported the attached claims.
| Authority tier | Citation occurrences | Occurrence share | Unique URLs | Unique-URL share |
|---|---|---|---|---|
| Tier A | 43 | 44.7917% | 21 | 42% |
| Tier B | 22 | 22.9167% | 13 | 26% |
| Tier C | 13 | 13.5417% | 9 | 18% |
| Tier D | 18 | 18.75% | 7 | 14% |
Weighted authority coverage used a different denominator and a stricter rule: only the best fully supporting adjacent source could earn authority credit for a required claim. On that measure, the run scored 67.5/132, or 51.1364%.
Failed hard gates
| Overall gate | Required | Audited result | Decision |
|---|---|---|---|
| Macro score | At least 90/100 | 61.3899 | FAIL |
| Batches passed | 6/6 | 0/6 | FAIL |
| Weighted factual accuracy | At least 95% | 87.5% | FAIL |
| Citation URL validity | 100% | 79.1667% | FAIL |
| Citation support precision | At least 95% | 69.8225% | FAIL |
| Citation completeness | At least 98% | 56.8182% | FAIL |
| Authority coverage | At least 90% | 51.1364% | FAIL |
| Fabricated or unresolvable citations | Zero | No fabricated URLs, but 20 invalid or unverifiable occurrences | FAIL |
| Critical-severity events | Zero | Zero | PASS |
| Fresh-chat contamination | Zero | Zero; the accidental duplicate was excluded | PASS |
Every batch also failed:
| Batch | Why the batch hard gate failed |
|---|---|
| B01 | Macro 67.5040, zero of four questions passed, and two invalid or unverifiable occurrences |
| B02 | Macro 43, zero of four questions passed, and four invalid or unverifiable occurrences |
| B03 | Macro 82.4583, two of four questions passed, and two invalid occurrences |
| B04 | Macro 76.9048 and only one of four questions passed, despite 100% URL validity |
| B05 | Macro 60, one of four questions passed, and eleven invalid or unverifiable occurrences |
| B06 | Macro 38.4722, zero of four questions passed, and one unverifiable occurrence |

Example of a fully supported citation

Q10 provides the cleanest positive example.
Question: What does the Consumer Price Index measure, which major groups are included, and are directly associated sales and excise taxes included?
DeepSeek claim: The CPI measures the average change over time in prices consumers pay for a representative basket of goods and services.
Cited source: Consumer Price Index Frequently Asked Questions – U.S. Bureau of Labor Statistics
Why it passed: The BLS page is the original agency source for the CPI definition. It directly supported the basket definition, the eight major groups, and the treatment of directly associated sales and excise taxes. All three Q10 occurrences were valid, Tier A, fully supporting, and adjacent to the relevant claims.
This example earned valid URL status, full support, Tier A authority, and precise placement. Q10 received 100 for factual accuracy, URL validity, expanded-edge support, completeness, authority, and placement.
A working link that did not support the whole claim
Q15 shows why a reachable page is not enough.
DeepSeek claim: NIST requires systems not to impose composition rules, and it requires screening passwords against a blocklist of commonly used, compromised, or weak passwords.
Cited source: NIST SP 800-63B Rev. 4 Password Updates
What the source established: The manually confirmed commercial article discussed password-rule updates and supported a separate citation occurrence attached to the 15-character single-factor minimum.
What it did not establish for the audited occurrence: The occurrence attached to the combined composition-rule and blocklist sentence received no support credit for either expanded edge. The answer also replaced the final standard’s named "expected" blocklist category with the looser word "weak."
The destination was reachable after manual confirmation, but the two attached edges received no support. The same page could support one claim in one place and fail another claim elsewhere.

Did DeepSeek prefer primary sources?
DeepSeek used a frozen preferred source or an exact official equivalent in 12 of 24 questions, or 50%: Q03, Q05, Q06, Q09, Q10, Q11, Q12, Q13, Q14, Q17, Q20, and Q21.
That diagnostic is not a pass rate. Q05, Q06, Q09, Q12, Q14, Q20, and Q21 still failed at least one hard gate. A source could also receive Tier A even if its URL differed from the frozen preferred page, provided it was an appropriate original authority and fully supported the claim.
At the citation-occurrence level, 43 of 96 citations were Tier A. At the stricter claim level, authority coverage was only 51.1364% because unsupported, incomplete, or low-authority citations earned reduced or zero credit.
For academic work, the DeepSeek research paper guide covers literature discovery, source verification, disclosure, and the boundary between finding a reference and citing evidence you have actually read.
What DeepSeek’s own terms say about Search accuracy
DeepSeek does not promise that Search makes every output correct. Its Terms of Use, updated March 27, 2026, state that:
- users who publish or disseminate outputs should verify authenticity and accuracy;
- outputs may include information from third-party websites or external sources;
- DeepSeek does not guarantee the validity, authenticity, legality, or security of external content;
- outputs may contain errors or omissions and are for reference;
- enabling Search may improve accuracy to some extent, but inaccuracy cannot be entirely avoided;
- outputs used for decisions with legal or material effects on people require human review.
Those terms do not prove that a particular citation is right or wrong. They explain why a visible source marker should start verification rather than end it.
The broader risk is not unique to one product. The NIST Generative AI Profile treats confidently presented false, misleading, or unsupported content as a risk requiring testing, evaluation, verification, and validation. NIST does not provide a DeepSeek score and is not cited as an endorsement of this benchmark.
Privacy note before using DeepSeek Search
DeepSeek’s Privacy Policy, updated February 10, 2026, says the company integrates third-party APIs for search and shares input keywords to provide search services. It also describes collection of prompts, uploaded files, chat history, device and network information, and feature-use logs. The policy says the services are not designed or intended to process sensitive personal data and asks users not to provide it.
Use public, non-sensitive search questions unless your organization has approved a different workflow. Do not put confidential client names, unreleased products, private incidents, medical details, personal records, trade secrets, or unpublished research into a hosted search prompt merely because the answer may contain citations.
Our DeepSeek privacy policy guide explains the hosted-data boundary in more detail. For governance and high-risk decisions beyond citation quality, see the guide to whether DeepSeek is safe.
How to verify a DeepSeek citation in six steps
- Open the exact cited URL. Do not verify a source from its title, citation label, or search snippet alone.
- Confirm the source identity. Check the publisher, author or agency, document title, and final destination after redirects.
- Check the date and version. Look for an effective date, update date, named revision, historical year, correction, or superseding notice.
- Find support for the exact claim. Split compound sentences and verify every number, unit, date, qualifier, exception, and comparison.
- Prefer the original record. Replace an aggregator or derivative summary with the law, standard, dataset, official announcement, paper, or agency page when available.
- Record uncertainty. Remove or replace a citation that cannot be verified. Do not convert a paywall, bot block, or temporary failure into a confident support judgment.
For an academic reference, verify its DOI and bibliographic metadata through the publisher, Crossref, or an appropriate scholarly database, then read the source. For a current rule or statistic, confirm the effective date and population definition. For market and competitor work, use a source-checked market research workflow rather than copying a generated answer directly into a decision document.
Prompt Template for More Verifiable DeepSeek Search Citations
The following prompt cannot guarantee correct citations, but it makes the evidence boundary explicit:
Use Search and answer only with claims you can support from sources you opened. Prefer original or official sources. Place a citation immediately after every factual sentence. For each source, provide the title, publisher, date, and direct URL. If a source does not explicitly support a clause, remove that clause. If evidence conflicts or is insufficient, state the conflict or uncertainty. Do not invent a title, author, date, DOI, quotation, statistic, or URL.
For a multi-part task, add:
Separate each independently checkable claim. Preserve units, denominators, dates, versions, exceptions, and scope. After drafting, create a claim-to-source table and mark any claim that lacks full support.
Treat the table as an audit queue, not proof. A model can incorrectly approve its own citation, so the final check must open the source outside the answer.
When DeepSeek Search is not enough
DeepSeek Search may help discover and summarize public sources. It is not a substitute for:
- a systematic or exhaustive literature search;
- a qualified legal, medical, financial, tax, security, or standards review;
- a licensed database needed to inspect controlling evidence;
- document version control and approval;
- privacy review for confidential or regulated material;
- source permissions and access control;
- independent human validation of a consequential claim.
If an organization needs answers grounded in an approved private corpus, a controlled retrieval system is a different problem. Our DeepSeek RAG knowledge base guide covers retrieval, permissions, source identifiers, evaluation, and auditability. It should not be confused with the hosted consumer Search feature tested here.
Students should follow institutional rules and cite evidence they verified, not a chatbot as a substitute for that evidence. The DeepSeek guide for students covers that broader workflow.
Protocol deviations
Three operator-side deviations were recorded across both protocols. In the July B02 baseline, a second identical prompt was accidentally submitted after the first response appeared absent in an accessibility snapshot. Only the untouched first completed response was frozen and scored. In the August expansion, the original US09 and IN11 captures ended before a complete answer was preserved. Each incomplete capture was excluded, and the identical prompt was rerun once in a fresh chat; the first completed rerun became the canonical scored answer.
| Batch | Deviation | Why it occurred | Effect on scoring or interpretation |
|---|---|---|---|
| B02 | A second identical prompt was accidentally submitted | The first response was not visible in an apparently empty accessibility snapshot | Only the untouched first response was frozen; the duplicate was excluded and was never counted or scored |
| US09 | The first capture ended before the answer completed | Operator-side evidence capture ended early | The incomplete capture was excluded; one identical-prompt rerun supplied the first complete answer and was scored |
| IN11 | The first capture ended before the answer completed | Operator-side evidence capture ended early | The incomplete capture was excluded; one identical-prompt rerun supplied the first complete answer and was scored |
The protocol allowed one clean rerun only when an operator-side capture failure prevented a complete answer from being preserved. A completed but weak, partially uncited, or inaccurate response remained the scored first response. No retry was allowed merely because an answer was poor.
In B04, screenshot capture timed out twice. That was an evidence-capture limitation, not a second scored-response deviation, because the live text, rendered HTML, and citation inventory were preserved.
Limitations of this citation audit
This was a controlled editorial benchmark, not a permanent product certification.
- It used two dated test periods in a signed-in consumer web session; the account plan and identifier were not published.
- Search results can vary by time, region, language, index state, account, and product update.
- Only first completed responses were scored; another generation could differ.
- The 48 English questions covered six global evidence domains plus U.S. and Indian official-source tasks, but they were not exhaustive.
- The benchmark emphasized public sources with independently verifiable ground truth.
- Paywalls, bot blocks, regional restrictions, and transient server errors can limit support adjudication.
- Atomic-claim segmentation, support labels, and authority tiers require reviewer judgment.
- For the July baseline, one primary reviewer completed the publication scoring. An independent checkpoint cross-checked all preserved baseline answers, required claims, citation occurrences, URL denominators, and the excluded B02 duplicate, but it did not complete a de novo page-level entailment rating. No V1 inter-rater agreement statistic is available.
- That independent checkpoint covered the July baseline. The August country expansion received a separate source-by-source adjudication and a consistency QA pass, but not a second de novo blind rater; no V2 inter-rater agreement statistic is available.
- The factual checkpoint initially differed from the primary ledger by one weighted point across five claims. Reconciliation retained 115.5/132 because the primary labels followed the frozen exact-qualifier and typical-use wording.
- The 169-edge reconciliation uses source-specific claim-level decisions for B01-B03 and repairs edge multiplicity while preserving occurrence-level support labels for B04-B06. A fully de novo second source-entailment review of B04-B06 could change individual edge labels.
- The independent checkpoint questioned the occurrence-level Tier A label assigned to a Washington State procurement PDF for a WCAG claim. Q17’s claim-level authority credit also had direct W3C support, so that flag did not change the question result.
- The test does not establish behavior for DeepSeek’s API, open-weight models, third-party hosts, no-search mode, uploaded private files, other accounts, other regions, or future versions.
- The U.S. and India labels identify the official-source subject matter. They do not prove the physical location of the browser, retrieval system, or model backend.
- The July and August protocols have different chat units and therefore retain separate macro scores; no combined 48-question score is reported.
- No independent country-subset PASS rule was preregistered; country score files label the subsets descriptively under inherited V1 gates while this article foregrounds question outcomes and exact metrics.
- The V2 prompts did not request an exact question-ID heading, but the inherited frozen rubric included one format point for it. Both subsets lost the component consistently without an additional severity deduction.
- The public release makes the benchmark rerunnable and the aggregate calculations reproducible, but it excludes private chat URLs, full generated answers, and raw interface evidence. Independent readers therefore cannot re-adjudicate every historical citation edge from the public package alone.
- Citation performance does not measure every aspect of answer usefulness, bias, safety, completeness, or reasoning quality.
The baseline preferred sources were verified on July 26, 2026, and the country-expansion sources were verified on August 6, 2026. Historical-year and named-document questions remained anchored to the date or version in the prompt. Source drift found during scoring was recorded rather than silently changing the gold standard.
Final verdict: can DeepSeek cite sources accurately?
DeepSeek can produce accurate, well-cited answers, but neither dated protocol was reliable enough for unverified research or consequential decisions.
The July baseline reached 87.5% weighted factual accuracy but only 56.8182% weighted citation completeness. The U.S. expansion passed 4 of 12 questions with 89.2473% factual accuracy and 75.2688% completeness. The India expansion passed 0 of 12, with 68.6869% factual accuracy, 34.3434% completeness, and one critical reversal of current UIDAI guidance. The safest use is source discovery followed by independent, claim-by-claim verification.
Do not generalize the result beyond the tested consumer Search configuration and date. A DeepSeek citation should remain a verifiable pointer to evidence, not a substitute for reading the evidence.
For more practical workflows, browse the DeepSeek use-case library.
Frequently asked questions
Can DeepSeek cite sources accurately?
DeepSeek can display and sometimes correctly use sources when Search is enabled, but the 48 tested questions did not show dependable unverified citation performance. The July baseline passed 4/24 questions, the U.S. expansion passed 4/12, and the India expansion passed 0/12. The protocols retain separate scores because their chat units differ. Treat the citations as leads that require verification.
Does DeepSeek use live web search?
DeepSeek’s official app announcement lists Web Search as a feature, and its Terms refer to a user-enabled Search function. Availability and behavior can vary by product surface, account, region, and time. This audit tests only the documented consumer Chat configurations used on July 26 and August 6, 2026.
Does enabling Search make every DeepSeek answer accurate?
No. DeepSeek’s Terms say Search may improve accuracy to some extent, while also stating that inaccurate generated content cannot be entirely avoided. Weighted factual accuracy was 87.5% in the July baseline, 89.2473% in the U.S. expansion, and 68.6869% in the India expansion. Far fewer questions cleared every answer and citation gate.
What is the difference between a valid link and an accurate citation?
A valid link reaches the intended substantive page. An accurate citation must also support the exact attached claim, use an appropriate source and version, and be placed clearly. The July baseline found 79.1667% URL validity but 69.8225% support precision across 169 claim edges.
Does a 200 OK response prove that a citation is correct?
No. 200 OK means the request succeeded. It does not prove that the page is the intended document, that it is authoritative or current, or that its content supports DeepSeek’s claim.
Can DeepSeek invent citations or URLs?
Generative systems can produce plausible but incorrect references. The July baseline identified no fabricated URL, but it still found 8 invalid and 12 unverifiable citation occurrences. The country expansion also found no fabricated URL, while keeping invalid and unverifiable states separate. Do not label a citation fabricated solely because an automated request failed; redirects, bot blocks, paywalls, and temporary errors require separate review.
Are DeepSeek citations reliable for academic research?
Treat them as discovery leads, not final references. Verify the title, authors, publication, year, DOI, and source record through the publisher or a scholarly index, then read the work to confirm it supports the claim. Follow institutional and publisher rules for AI use and disclosure.
Does the DeepSeek API automatically provide web citations?
Do not infer API behavior from the consumer Chat Search interface. Official API documentation and the tested endpoint must establish any retrieval or citation feature. This audit does not score the API, open-weight models, or third-party DeepSeek hosts.
How do I verify a DeepSeek source?
Open the exact URL, confirm source identity, check the date and version, locate support for every adjacent clause, and prefer the original record when available. Remove or replace a source that cannot be verified, and record uncertainty instead of guessing.
What metrics should a citation audit report?
A transparent audit should report URL resolution, target identity, claim-level support, citation completeness, source authority, freshness, placement clarity, and answer correctness. Each rate needs a published denominator so readers can distinguish citation occurrences, citation-claim edges, weighted claims, questions, and batches.
How should redirects, paywalls, and blocked pages be scored?
Score them as separate states. A relevant redirect can pass after its destination is checked. A paywalled or bot-blocked source may be identifiable but not fully verifiable. A dead page, irrelevant redirect, or deceptive destination is a different failure. Automated inaccessibility alone does not prove fabrication.
Are primary sources always better than secondary sources?
Primary sources are usually preferable for official rules, specifications, statistics, court decisions, research findings, and company announcements. Strong secondary reporting can add context or challenge a primary claim. The right test is whether the source is authoritative, current, and appropriate for the exact statement.
Can one citation support several claims in the same sentence?
Yes, but every associated clause must be tested. The July baseline expanded 96 citation occurrences into 169 citation-claim edges because compound sentences often contained several atomic claims. One page may support all, some, or none of them.
How can I prompt DeepSeek to provide more verifiable sources?
Tell it to use Search, prefer primary sources, place a citation immediately after each factual sentence, provide direct URLs and dates, and state when evidence is insufficient. Instruct it not to invent titles, authors, quotations, DOIs, statistics, or links. Better prompting improves traceability but does not replace human verification.
Should I cite DeepSeek or the original source?
For factual claims, cite the original source you independently verified whenever possible. Cite or disclose DeepSeek separately when a style guide, institution, publisher, employer, or research protocol requires disclosure of AI assistance. Do not use the chatbot as a substitute for the evidence it points to.
