Can DeepSeek Cite Sources Accurately? A Reproducible 48-Question Audit

A reproducible 48-question audit of DeepSeek Search citations, including separate U.S. and India extensions with downloadable data and claim-level scoring.

Can DeepSeek cite sources accurately? Across 48 English questions in 30 canonical scored chats, our answer was not reliable enough for unverified use. The research combines a July 2026 global-domain baseline with a separately scored August 2026 country expansion covering official sources from the United States and India. Two additional incomplete capture attempts were excluded under the documented rerun rule.

In the July baseline, 76 of 96 citation occurrences resolved to valid substantive sources or were confirmed manually, for 79.1667% URL validity. The stricter claim-level review found 118 supported-edge points across 169 citation-claim edges, or 69.8225% support precision. Weighted citation completeness was 75/132, or 56.8182%. DeepSeek passed 4 of 24 questions and 0 of six batches, producing a final macro score of 61.3899/100.

Those figures measure different things. A valid URL can lead to the wrong evidence. A relevant page can fail to support the sentence beside it. A citation can support one clause while leaving another ungrounded. An otherwise correct answer can also omit a material date, denominator, exception, version, or scope qualifier.

Benchmark scope: We tested the official DeepSeek consumer web chat with Search on and DeepThink off. The July baseline used six fresh chats, each containing one exact four-question batch. The August country expansion used 24 additional fresh chats, one exact question per chat. Only the first completed answer was scored, except for two documented operator-side capture failures that were excluded and rerun once under the same prompt. The results do not establish the behavior of the DeepSeek API, open-weight models, third-party hosts, no-search chat, other accounts, other regions, or future product versions.

For broader product context, start with our independent DeepSeek guide. This article stays deliberately narrow: it measures whether citations returned by two dated DeepSeek Search protocols were reachable, correctly placed, authoritative, complete, and supportive. The complete question packs, results, methodology, and machine-readable files are available in the public DeepSeek citation-audit repository.

DeepSeek citation accuracy: the short answer

The July baseline, deepseek-web-citation-audit-v1 version 1.0.0, used 24 English questions covering health guidance; Earth, space, and climate science; public data and economics; law, regulation, and security policy; web and software standards; and history and world heritage.

The overall result was FAIL under the benchmark’s predeclared hard gates.

MeasureAudited resultDenominatorWhat it means
Overall macro score61.3899/10024 question scoresMean of the final question-level numeric scores
Questions passed4/2424 questionsQuestions meeting the numeric threshold and every question hard gate
Batches passed0/6Six four-question batchesBatches meeting the numeric threshold and every batch hard gate
Weighted factual accuracy115.5/132 (87.5%)132 gold-claim weight pointsWhether required facts, numbers, dates, scope, and qualifiers were correct
Citation URL validity76/96 (79.1667%)96 citation occurrencesWhether each displayed citation reached the intended substantive source
Citation support precision118/169 (69.8225%)169 citation-claim edgesWhether cited sources entailed the atomic claims a reader would associate with them
Weighted citation completeness75/132 (56.8182%)132 gold-claim weight pointsWhether correct required claims had adjacent, valid, fully supporting citations
Weighted authority coverage67.5/132 (51.1364%)132 gold-claim weight pointsWhether full support came from sources appropriate to the claim
Citation placement precision70/96 (72.9167%)96 citation occurrencesWhether citations were adjacent to at least one supported claim
Preferred-source hit rate12/24 (50%)24 questionsQuestions using a frozen preferred source or an exact official equivalent

Most important strength: Weighted factual accuracy reached 87.5%, and Q10, Q11, Q13, and Q17 combined accurate answers with complete and authoritative support.

Most consequential failure: Only 56.8182% of required claim weight received full citation-completeness credit. Four answers had no rendered citations, and two more had citations but no URL that could earn validity credit.

Practical recommendation: Use DeepSeek Search to discover possible sources, not as a final authority. Open the exact URL and verify every material clause before relying on, publishing, or citing the answer.

August 2026 U.S. and India country expansion

On August 6, 2026, we added 24 new questions: 12 tied to verified U.S. official sources and 12 tied to verified Indian official sources. Each canonical result used its own fresh chat under the same visible consumer settings: Instant mode, Search on, DeepThink off, English only, and the first completed answer scored. Two incomplete operator-side captures were excluded and rerun once as documented below.

We report this expansion separately from the July baseline. The baseline submitted four questions per chat and used a full batch gate; the expansion submitted one question per chat and has no batch score. Combining their macro scores would hide that protocol difference.

Country subsetQuestions passedMacro scoreFactual accuracyURL validitySupport precisionCompletenessAuthority coveragePlacement
United States4/1274.8202/10083/93 (89.2473%)103/115 (89.5652%)147.5/180 (81.9444%)70/93 (75.2688%)61/93 (65.5914%)102/115 (88.6957%)
India0/1242.1966/10068/99 (68.6869%)61/80 (76.25%)91.5/129 (70.9302%)34/99 (34.3434%)25/99 (25.2525%)60/80 (75%)

No distinct country-subset pass rule was frozen before the live run. The canonical outcomes are the question-level pass counts and the metrics above. The machine-readable score files also provide a descriptive FAIL label by applying the inherited V1 hard-gate logic, but it should not be read as a separately preregistered country-aggregate outcome.

The U.S. subset passed four questions: constitutional qualifications, post-1977 copyright duration, CISA Secure by Design principles, and USGS earthquake terminology. The subset missed the inherited aggregate thresholds. One DOT answer omitted material refund conditions; the EEOC answer replaced the special age-discrimination filing condition; and the FEMA answer misstated or overgeneralized waiting-period exceptions. Several other answers were factually strong but failed support, authority, or unresolvable-link gates.

The India subset passed no question. The most serious error appeared in the UIDAI e-KYC answer: it said biometric authentication was prohibited, while current UIDAI guidance says e-KYC authentication is performed using OTP and/or biometric authentication. Other material failures included reversing the cancellation-charge condition in the E-Commerce Rules, reporting non-final Census 2011 population figures instead of the frozen final values, and omitting required probabilities and error bounds from the dated April 13, 2026 IMD monsoon forecast.

These country labels describe the official-source subject matter, not the physical location of the user, browser, retrieval service, or model backend. The consumer interface did not expose those systems, so this is not a latency or geographic-routing benchmark.

Bar chart comparing DeepSeek citation accuracy metrics for U.S. and India official-source question sets.
DeepSeek citation accuracy on separate 12-question U.S. and India official-source subsets, tested August 6, 2026. This is not a geographic routing test.

What changed in version 2.0.0

  • One question was submitted per fresh chat to isolate context and make question-level reruns auditable.
  • The U.S. pack froze 12 questions, 93 weighted claim points, and official sources from bodies including Congress.gov, the Copyright Office, EEOC, DOT, NIST, CISA, FDA, USDA, TreasuryDirect, USGS, and FEMA.
  • The India pack froze 12 questions, 99 weighted claim points, and official sources including India Code, Consumer Affairs, DICGC, UIDAI, ECI, ISRO, Census of India, IMD, BIS, FSSAI, and MeitY.
  • The same dimensions remained separate: factual accuracy, citation URL validity, claim-level support, completeness, authority, placement, and format.
  • The inherited format rubric awarded one point for an exact question-ID heading even though the V2 prompts did not request one. Both subsets lost that one component consistently, but no separate severity deduction was added for an instruction the prompts never gave. The mismatch did not cause any otherwise passing question to fail.

Download the data and reproduce the audit

The public GitHub repository contains the frozen question packs, gold claims, methodology, scored results, data dictionaries, checksums, and a CC BY 4.0 license for the original release data and documentation.

Private chat URLs, account identifiers, and raw interface snapshots are intentionally excluded. The public files preserve prompts, gold claims, official source records, scores, exact denominators, findings, input hashes, and protocol notes without exposing the signed-in session.

What counts as an accurate DeepSeek citation?

Citation accuracy is not one binary property. We used seven checks because each catches a different failure.

DimensionPass conditionWhy it matters
URL validityThe citation reaches the intended substantive source within five redirects, or a normal browser confirms a blocked official sourceA broken, deceptive, or unresolvable link cannot be verified
Target identityThe final page is the source described by the citation labelA real URL can still point to the wrong article, document, or publisher
Claim supportThe source directly entails the adjacent atomic claimA topically related page may not prove the model’s statement
CompletenessEvery correct required claim has an adjacent valid citation with full supportA few good citations do not ground an entire answer
Source authorityThe source is appropriate for the exact claimAn original law, standard, dataset, or agency record is stronger than a derivative summary within its remit
Freshness and versionThe source matches the date, revision, or historical frame requestedA formerly correct page can be wrong for a current-rule question
PlacementA reasonable reader can identify which claim the citation supportsA detached or decorative citation makes claim-level verification difficult

The distinction between answer correctness and citation quality follows the core logic of the ALCE citation-evaluation benchmark. We also split compound sentences into independently checkable claims, reflecting the fine-grained citation problem examined by ALiiCE. Our practical method is documented under our editorial and source standards. It is an independent adaptation for live consumer web citations, not an official ALCE or ALiiCE replication.

A valid page is not automatically valid evidence

Suppose an answer says:

Agency X adopted Rule Y in 2025, and the rule requires Z.

One source might confirm the agency and adoption date but describe a different requirement. The URL can return a successful page, the publisher can be authoritative, and the citation can still fail the requirement claim.

The same boundary applies to HTTP status checks. Under RFC 9110, 200 OK indicates that the request succeeded. It does not establish that the returned page is the intended document, that the information is current, or that the page supports DeepSeek’s adjacent statement.

How the final edge denominator was established

The first scoring pass created one nearest-claim support judgment per citation occurrence. That was useful as a conservative checkpoint, but 96 occurrences could not serve as the final claim-edge denominator because the frozen rubric requires an edge for every atomic claim that a reasonable reader would associate with a citation.

The reconciliation expanded compound citation clusters to 169 citation-claim edges. It included gold-aligned atomic response claims plus two unambiguous extra factual claims, and it split Q06’s river and fresh-lake percentages because either number could be true without the other. B01-B03 received source-specific edge labels; B04-B06 used the expanded atomic edge map while preserving the primary support label at each citation occurrence. The resulting 118/169 figure is the canonical conservative reconciliation for this publication. A de novo second page-level review of B04-B06 could change individual labels, but it could not justify returning to a one-edge-per-occurrence denominator.

The source-specific review changed four earlier inherited-label decisions. It treated the Q05 Horizons launch-date edge as partial and its phase-name edge as unsupported; found that the Q06 USGS table contradicted the river percentage while supporting the fresh-lake percentage; credited the Q09 Census source for both resident-population exclusion and home-state apportionment treatment; and denied Q12 support for an active-duty permanent-change-of-station scope that its cited rate document did not establish.

How we designed the July 24-question citation benchmark

The questions, prompts, gold claims, source preferences, scoring formulas, thresholds, and hard gates were frozen before the first live submission. Changing them after seeing an output would create a new benchmark version.

The frozen design contained:

  • 24 scored questions;
  • six topic batches;
  • four questions per batch;
  • 101 required atomic claims;
  • 132 total required-claim weight points;
  • 28 preferred Tier A source records;
  • one exact prompt per batch;
  • one fresh chat per batch;
  • Search on and DeepThink off;
  • first completed responses only;
  • no regeneration, repair prompt, clarification, or substantive follow-up.

The 28 preferred records were original agencies, standards bodies, laws, official datasets, mission records, treaty text, or institutional primary sources. An equivalent official source could receive the same Tier A rating even when its URL differed from the frozen preferred page.

Infographic showing six independent DeepSeek chats, four questions per batch, and the frozen test conditions.
The audit used six fresh chats, four questions per batch, Search on, DeepThink off, no regeneration, and the first completed response only.

Six batches, six evidence domains

BatchQuestionsDomainExamples of required evidence
B01Q01-Q04Health guidanceCDC activity and sleep guidance, FDA caffeine guidance, NIH vitamin D guidance
B02Q05-Q08Earth, space, and climate scienceNASA mission records, USGS water estimates, ozone authorities, IPCC findings
B03Q09-Q12Public data and economicsCensus definitions, BLS CPI scope, Federal Reserve framework, historical IRS rates
B04Q13-Q16Law, regulation, and security policyFair-use factors, CAN-SPAM requirements, final NIST password rules, GDPR Article 17
B05Q17-Q20Web and software standardsWCAG contrast, HTTP 429, robots exclusion, official Python semantics
B06Q21-Q24History and world heritageLibrary of Congress, National Archives, UNESCO, and National Park Service records

The questions were designed to expose recurring citation problems: numbers without units or denominators, obsolete revisions, partly supported compound sentences, legal rights without exceptions, statistics without population definitions, and technically correct claims sourced to weak derivative pages.

Exact response contract

Every batch prompt required DeepSeek to:

  1. answer all four questions;
  2. use the exact question ID as an H2 heading;
  3. stay within each question’s sentence range;
  4. cite every factual sentence inline with clickable citations;
  5. prefer original agencies, official standards, treaty text, mission records, and institutional primary sources;
  6. avoid search-results pages, AI-generated pages, citation aggregators, and unsourced snippets;
  7. add no preface, bibliography, conclusion, or advice outside the four answers.

The experiment’s unit was one first completed response to one four-question batch prompt. Every batch started in a fresh chat, so previous answers could not become context for later topics.

Controlled interface conditions

SettingRecorded value
Product surfaceOfficial DeepSeek consumer web chat
Test dateJuly 26, 2026
Test window11:27:38 a.m. to 11:41:21 a.m. PDT (18:27:38 to 18:41:21 UTC)
Visible mode labelInstant
SearchOn for all six batches
DeepThinkOff for all six batches
LanguageEnglish only
Account scopeOne signed-in consumer web session; account plan and identifier were not exposed or published
ChatsSix fresh chats
Regenerations0
Corrective follow-ups0
Recorded protocol deviations1

The consumer interface did not expose the underlying model identifier or version, backend routing, retrieval index or provider configuration, account plan, region, browser, or interface locale. We do not infer those values. A browser test should not be presented as a benchmark of a named API model unless the interface establishes that identity.

DeepSeek’s official app announcement lists Web Search among the product’s features. That documentation does not guarantee identical retrieval behavior across every interface, account, region, or date.

Consumer Search and API retrieval are separate product questions. Our DeepSeek API Web Search guide documents that boundary; this audit does not infer API capabilities from buttons visible in the chat interface.

How every answer and citation was scored

Flow diagram from required claim through citation marker, resolved URL, source content, and final support and authority judgment.
The audit treats factual accuracy, URL validity, claim support, citation completeness, placement, and source authority as separate questions.

Scoring began after all six first responses and their citation inventories were frozen. Each answer was split into atomic claims before citation quality was judged.

SectionMaximum pointsWhat it measured
Factual accuracy35Required claims against the frozen gold answer
URL validity15Citation occurrences reaching intended substantive sources
Claim support20Source entailment across citation-claim edges
Citation completeness15Correct required claims with full adjacent support
Source authority10Best fully supporting source tier for each required claim
Format and placement5Exact heading, length, prohibited extras, and placement precision

A question needed at least 85/100, at least 90% factual accuracy, 100% URL validity, at least 90% support precision, 100% completeness, at least 80% authority coverage, no fabricated or unresolvable citation, no critical event, and no material contradiction of a required claim.

A batch needed at least 90/100, all four questions to pass, no invalid or unverifiable citation occurrence, no critical event, and all exact question headings. The full benchmark needed at least 90/100, all six batches to pass, at least 95% overall factual accuracy, 100% URL validity, at least 95% support precision, at least 98% completeness, at least 90% authority coverage, no fabricated or unresolvable citation, no critical event, and no contamination between fresh chats.

The numeric score did not override hard gates. That is why Q03 scored 96.3333 but failed its 90% support gate, and why Q13 scored 97 yet still passed despite a small format deduction.

Teams adapting this method to their own applications can use the DeepSeek evaluation framework to separate datasets, evaluators, thresholds, and regression runs. The benchmark files in this research release remain the source of truth for the numbers reported here.

URL states were kept separate

  • Valid: the citation resolved to the intended substantive page or official document.
  • Valid manual: automated access was blocked, but independent retrieval confirmed the intended publisher, page, and relevant content.
  • Invalid: the link was malformed, broken, deceptive, unrelated, a soft 404, a search-results page, or a source-free internal citation page.
  • Unverifiable: neither automated nor manual inspection could establish the target and substantive content.

Redirect chains were followed for no more than five redirects. Raw links were preserved, while normalized links were used only for duplicate detection. Automated inaccessibility alone was not labeled fabrication.

Source authority was scored only after support

TierSource typeScore
AOriginal law, regulation, standard, official dataset, mission record, or authoritative agency page1.00
BHigh-quality institutional synthesis or scholarly secondary source0.75
CGeneral secondary source, commercial explainer, news report, or encyclopedia0.25
DUser-generated, AI-generated, scraped, low-accountability, citation-aggregation, or search-results source0.00

A prestigious domain did not earn automatic authority credit. The page first had to support the claim.

Results across all six citation batches

Scorecard for the 24-question DeepSeek citation audit showing the final hard-gate failure and six measured metrics.
The frozen audit scored 61.39/100: four of 24 questions passed, no batch passed, and claim support measured 118/169 under the conservative edge method.
BatchDomainMacro scoreQuestions passedFactual accuracyURL validitySupport precisionPlacementCompletenessAuthority coverageResult
B01Health guidance67.50400/483.3333%88.8889%16.5/26 (63.4615%)15/18 (83.3333%)72.2222%52.7778%FAIL
B02Earth, space, and climate science430/469.0476%75%14.5/25 (58%)11/16 (68.75%)42.8571%42.8571%FAIL
B03Public data and economics82.45832/491.3043%85.7143%18/23 (78.2609%)11/14 (78.5714%)82.6087%82.6087%FAIL
B04Law, regulation, and security policy76.90481/494.4444%100%45/51 (88.2353%)20/21 (95.2381%)81.4815%73.1481%FAIL
B05Web and software standards601/494.7368%47.619%20/35 (57.1429%)10/21 (47.619%)31.5789%22.3684%FAIL
B06History and world heritage38.47220/489.5833%83.3333%4/9 (44.4444%)3/6 (50%)25%25%FAIL

B04 is the clearest warning against using link reachability as a proxy for answer quality: all 21 citation occurrences were URL-valid, yet the batch still failed with one passing question, 81.4815% completeness, and 73.1481% authority coverage. B05 shows the opposite pattern in the same run: factual accuracy was 94.7368%, but only 10 of 21 citation occurrences were valid and the batch’s completeness was 31.5789%.

Stacked bars showing valid, manually verified, invalid, and unverifiable citation occurrences for six DeepSeek test batches.
URL validity ranged from 47.6% in B05 to 100% in B04; these rates measure destination validity, not whether each source supports the associated claim.

All 24 question results

The table reports final reconciled scores and hard-gate results separately. A high score cannot turn a failed hard gate into a pass.

QuestionTopicFinal scoreResultMost important audited finding
Q01Adult physical activity62FAILThe activity totals and smaller-chunk advice were right, but the all-major-muscle-groups qualifier was missing and one citation was unverifiable.
Q02Caffeine38.6825FAILThe 400 mg framing was preserved, but the coffee-cup conversion and caution population were materially incomplete.
Q03Vitamin D96.3333FAILAll NIH intake values were accurate and complete; one redundant Mayo citation only partially supported the combined age-band sentence, leaving support precision at 80%.
Q04Adult sleep73FAILThe three sleep bands were correct, but one of two citations was unverifiable and the surviving source was a general secondary explainer.
Q05Voyager 154FAILLaunch, targets, and August 2012 were right, but the required NASA distinction was replaced with a mission phase name; the Horizons date edge was only partial and its phase-name edge was unsupported.
Q06Earth’s water44FAILThe large reservoirs were accurate; the tiny rivers-and-lakes fraction was materially misstated, the USGS source contradicted the river value, and half the citation occurrences were broken.
Q07Ozone53FAILOzone composition and depletion chemistry were good, but the regional recovery timetable did not match the frozen benchmark.
Q08IPCC warming findings21FAILMost temperature values were right, but the central attribution wording was weakened and the answer contained no live citation.
Q092020 Census population78.3333FAILThe resident count, date, scope, and growth were correct, and expanded source support was 100%; the answer’s overseas home-state allocation wording remained factually incomplete.
Q10Consumer Price Index100PASSThe CPI basket, substitution distinction, and non-cost-of-living qualifier matched the BLS benchmark with official citations.
Q11Federal Reserve inflation goal100PASSThe longer-run 2% PCE framing and measure distinction were accurate, complete, and supported by Federal Reserve sources.
Q122024 IRS mileage rates51.5FAILAll three mileage rates were right, but two KPMG links were dead and the cited rate document did not establish the military permanent-change-of-station qualifier.
Q13Fair use97PASSAll four statutory factors and the case-by-case rule were accurate and supported; only the requested sentence range was missed.
Q14CAN-SPAM85FAILCAN-SPAM duties and timing were largely correct, but the advertising-identification exception was omitted.
Q15NIST password rules47.3333FAILSeveral password rules were directionally right, but citation support was weak and the named final-revision requirement was not met cleanly.
Q16GDPR erasure78.2857FAILThe grounds and exception categories were mostly accurate, but a legally material necessity limit and the requested sentence range were missed; reconciled support was 88.0952%.
Q17WCAG contrast100PASSContrast ratios, large-text thresholds, and exemptions were accurate and fully cited; the official W3C link used a WCAG 2.1 understanding path whose relevant ratios match WCAG 2.2.
Q18HTTP 42938FAILThe semantics were accurate, but every Stack Overflow revision citation remained unverifiable, so none earned support or completeness credit.
Q19Robots Exclusion Protocol38FAILThe robots.txt statements matched RFC 9309, but all six citation occurrences were unverifiable or broken.
Q20Python lists and tuples64FAILMutability and singleton syntax were right, but typical access patterns were incomplete and the supporting sources had low authority.
Q21Jamestown84.8889FAILAll required facts were accurate and an official Library of Congress source covered them, but weak placement and one unverifiable news citation reduced precision.
Q2219th Amendment38FAILAll five required facts matched the benchmark, yet the answer supplied zero citations.
Q23Great Barrier Reef9FAILYear, criteria, and area were right, but reef and island counts were wrong or over-precise and no citation was supplied.
Q24Yellowstone22FAILThe law, land withdrawal, and 1916 comparison were right; the National Park Service attribution was weakened and the entire answer was uncited.

The four passing questions were Q10, Q11, Q13, and Q17. In the final reconciliation, Q09’s support gate changed from fail to pass and Q16’s changed from pass to fail; both questions still failed other frozen gates. Q03 illustrates why the gates matter: its factual accuracy, URL validity, completeness, authority coverage, and placement were all 100%, but its expanded claim-edge support precision was 80%, below the required 90%.

Grid of all 24 DeepSeek citation-audit question scores with four passes and hard-gate failures marked.
Only Q10, Q11, Q13, and Q17 cleared the numeric threshold and every question-level hard gate.

Link failures, unsupported claims, and missing citation coverage

These failure classes remain separate. A broken link is not the same as an unsupported claim, and a missing citation is not the same as an incorrect fact.

Invalid or unverifiable citation occurrences

The ten grouped rows below account for all 20 non-valid occurrences: eight invalid and twelve unverifiable. Repeated citation IDs are retained because the benchmark denominator counts every occurrence.

Citation ID(s)QuestionDisplayed labelPublic targetValidityOccurrencesAudit note
B01-CIT05Q015Emanate Health GUID pageUnverifiable1The server returned 403, and independent retrieval did not establish the exact page title or substantive content.
B01-CIT17Q044DMU Clinic GUID pageUnverifiable1The server returned 403, and independent retrieval did not establish the exact page title or substantive content.
B02-CIT01Q051Archived Voyager 1 pageUnverifiable1A web-application challenge prevented inspection of the archived snapshot.
B02-CIT07, B02-CIT09, B02-CIT11Q0610Water distribution PDF targetInvalid3The target returned 404 with no substantive document.
B03-CIT11, B03-CIT13Q1215KPMG mileage PDF targetInvalid2The target returned a 404 page instead of the intended PDF.
B05-CIT07, B05-CIT09, B05-CIT11Q182Stack Overflow revision-source target AUnverifiable3A 403 challenge and cache miss prevented confirmation of the exact revision content.
B05-CIT08, B05-CIT10Q186Stack Overflow revision-source target BUnverifiable2A 403 challenge and cache miss prevented confirmation of the exact revision content.
B05-CIT12, B05-CIT14, B05-CIT16Q197IETF development-host RFC 9309 targetUnverifiable3A 403 challenge prevented confirmation of the exact development-host document.
B05-CIT13, B05-CIT15, B05-CIT17Q1910Third-party RFC 9309 PDF targetInvalid3The target returned 404 with no substantive RFC PDF.
B06-CIT03Q2113Politico Jamestown pageUnverifiable1A 403 challenge and robots restriction prevented confirmation of the exact page content.

No fabricated URL was identified. That does not rescue the run: invalid and unverifiable citations still failed the benchmark’s hard gates.

Reachable citations with no support credit

The final edge inventory contained 110 full-support edges, 16 partial-support edges, and 43 no-support edges. The full and partial edges yield 118 support points across 169 edges. Of the 43 no-support edges, 32 came from invalid or unverifiable occurrences listed above. The remaining 11 no-support edges came from nine citation occurrences whose targets were valid or manually confirmed:

Citation ID(s)QuestionNo-support edgesSourceAudited reason
B01-CIT11, B01-CIT12Q022Mayo Clinic caffeine articleThe manually confirmed page did not establish the displayed three-to-four-cup conversion or the separate sensitivity statement.
B02-CIT05Q051NASA JPL Horizons recordThe record described interstellar-mission timing but did not establish the displayed current-phase designation.
B02-CIT10Q061USGS Earth water distributionThe source contradicted the displayed 0.006% river value; a separate edge from the same occurrence fully supported the fresh-lake value.
B02-CIT15Q071Cambridge ozone research paperThe paper discussed signal timing but did not establish return to 1980 near-global total-column values under continued policy.
B03-CIT14Q121Taxpayer Advocate rate documentThe document listed the rate but did not establish the active-duty, military-orders, permanent-change-of-station scope.
B04-CIT14Q152NIST SP 800-63B Rev. 4 Password UpdatesThe manually confirmed commercial explainer received no support credit for the attached no-composition-rules and blocklist edges.
B06-CIT02Q212Jamestown classroom materialsAt this rendered position, the occurrence received no support credit for the establishment-date and charter edges; separate Library of Congress citations supplied the answer’s full claim coverage.
B06-CIT05Q211Jamestown classroom materialsAt this rendered position, the occurrence received no support credit for the first-permanent-settlement edge; another adjacent Library of Congress citation supplied coverage.

B01-B03 use source-specific edge evidence. B04-B06 preserve the primary occurrence-level support decisions while expanding edge multiplicity. A de novo second page-level review could revise individual B04-B06 labels, but it would not justify returning to one edge per citation occurrence.

Required claim weight without full completeness credit

Overall, 75 of 132 required-claim weight points received full completeness credit. The remaining 57 points were lost because the answer was wrong or incomplete, no adjacent citation appeared, a citation was invalid or unverifiable, or the adjacent source did not fully support the claim. The table lists every question below 100% completeness.

QuestionCredited weightCompletenessMain reason for the gap
Q013/560%The strength statement omitted the all-major-muscle-groups qualifier, and one associated citation was unverifiable.
Q022/540%The cup conversion was wrong and the caution scope omitted material populations and conditions.
Q053/560%The answer omitted the required first-human-made and most-distant-object distinction.
Q063/560%The nested rivers-and-lakes fraction was misstated, and three occurrences pointed to a dead PDF.
Q073/650%The regional recovery years did not match the frozen benchmark.
Q080/50%The attribution language was weakened and the answer rendered no citation.
Q095/771.4286%The home-state allocation of eligible overseas counts was incomplete.
Q123/560%Two PDF citations were dead and the military moving exception lacked its permanent-change-of-station condition.
Q147/887.5%The answer omitted the statutory exception to the advertising-identification duty.
Q153/650%The answer softened final-standard wording and relied on weak or obsolete secondary material.
Q166/785.7143%The research and archiving exception omitted the impossible-or-seriously-impaired purpose condition.
Q180/50%All five citation occurrences were unverifiable.
Q190/60%All six citation occurrences were invalid or unverifiable.
Q202/450%Required iteration, unpacking, or indexing details were omitted, and source authority was low.
Q220/60%The factual answer rendered no citation.
Q230/60%Reef and island counts were wrong or over-precise, and the answer rendered no citation.
Q240/60%The National Park Service attribution was weakened, and the answer rendered no citation.

Q03, Q04, Q10, Q11, Q13, Q17, and Q21 received 100% completeness credit. Completeness alone did not make all of them pass: URL validity, support precision, authority, placement, format, and the other hard gates still applied.

Source-tier distribution

The occurrence denominator is 96; the unique-URL denominator is 50. These counts describe the class of cited sources, not whether those sources supported the attached claims.

Authority tierCitation occurrencesOccurrence shareUnique URLsUnique-URL share
Tier A4344.7917%2142%
Tier B2222.9167%1326%
Tier C1313.5417%918%
Tier D1818.75%714%

Weighted authority coverage used a different denominator and a stricter rule: only the best fully supporting adjacent source could earn authority credit for a required claim. On that measure, the run scored 67.5/132, or 51.1364%.

Failed hard gates

Overall gateRequiredAudited resultDecision
Macro scoreAt least 90/10061.3899FAIL
Batches passed6/60/6FAIL
Weighted factual accuracyAt least 95%87.5%FAIL
Citation URL validity100%79.1667%FAIL
Citation support precisionAt least 95%69.8225%FAIL
Citation completenessAt least 98%56.8182%FAIL
Authority coverageAt least 90%51.1364%FAIL
Fabricated or unresolvable citationsZeroNo fabricated URLs, but 20 invalid or unverifiable occurrencesFAIL
Critical-severity eventsZeroZeroPASS
Fresh-chat contaminationZeroZero; the accidental duplicate was excludedPASS

Every batch also failed:

BatchWhy the batch hard gate failed
B01Macro 67.5040, zero of four questions passed, and two invalid or unverifiable occurrences
B02Macro 43, zero of four questions passed, and four invalid or unverifiable occurrences
B03Macro 82.4583, two of four questions passed, and two invalid occurrences
B04Macro 76.9048 and only one of four questions passed, despite 100% URL validity
B05Macro 60, one of four questions passed, and eleven invalid or unverifiable occurrences
B06Macro 38.4722, zero of four questions passed, and one unverifiable occurrence
Bar chart comparing source-tier distributions across 96 citation occurrences and 50 unique URLs in the DeepSeek audit.
Tier A sources accounted for 43 of 96 citation occurrences and 21 of 50 unique URLs; authority tier alone does not prove that a source supports a claim.

Example of a fully supported citation

DeepSeek Search response showing Q17 and Q18 with complete answer text and inline numbered citation markers.
The frozen B05 first response displayed inline citation markers for the WCAG and HTTP questions while Search was enabled and DeepThink was off.

Q10 provides the cleanest positive example.

Question: What does the Consumer Price Index measure, which major groups are included, and are directly associated sales and excise taxes included?

DeepSeek claim: The CPI measures the average change over time in prices consumers pay for a representative basket of goods and services.

Cited source: Consumer Price Index Frequently Asked Questions – U.S. Bureau of Labor Statistics

Why it passed: The BLS page is the original agency source for the CPI definition. It directly supported the basket definition, the eight major groups, and the treatment of directly associated sales and excise taxes. All three Q10 occurrences were valid, Tier A, fully supporting, and adjacent to the relevant claims.

This example earned valid URL status, full support, Tier A authority, and precise placement. Q10 received 100 for factual accuracy, URL validity, expanded-edge support, completeness, authority, and placement.

A working link that did not support the whole claim

Q15 shows why a reachable page is not enough.

DeepSeek claim: NIST requires systems not to impose composition rules, and it requires screening passwords against a blocklist of commonly used, compromised, or weak passwords.

Cited source: NIST SP 800-63B Rev. 4 Password Updates

What the source established: The manually confirmed commercial article discussed password-rule updates and supported a separate citation occurrence attached to the 15-character single-factor minimum.

What it did not establish for the audited occurrence: The occurrence attached to the combined composition-rule and blocklist sentence received no support credit for either expanded edge. The answer also replaced the final standard’s named "expected" blocklist category with the looser word "weak."

The destination was reachable after manual confirmation, but the two attached edges received no support. The same page could support one claim in one place and fail another claim elsewhere.

DeepSeek Search response showing Q22 and Q23 without inline citation markers in the captured answer sections.
Q22 and Q23 appeared without inline citation markers in the frozen B06 first response; the preserved B06 citation inventory contains citation occurrences only for Q21.

Did DeepSeek prefer primary sources?

DeepSeek used a frozen preferred source or an exact official equivalent in 12 of 24 questions, or 50%: Q03, Q05, Q06, Q09, Q10, Q11, Q12, Q13, Q14, Q17, Q20, and Q21.

That diagnostic is not a pass rate. Q05, Q06, Q09, Q12, Q14, Q20, and Q21 still failed at least one hard gate. A source could also receive Tier A even if its URL differed from the frozen preferred page, provided it was an appropriate original authority and fully supported the claim.

At the citation-occurrence level, 43 of 96 citations were Tier A. At the stricter claim level, authority coverage was only 51.1364% because unsupported, incomplete, or low-authority citations earned reduced or zero credit.

For academic work, the DeepSeek research paper guide covers literature discovery, source verification, disclosure, and the boundary between finding a reference and citing evidence you have actually read.

What DeepSeek’s own terms say about Search accuracy

DeepSeek does not promise that Search makes every output correct. Its Terms of Use, updated March 27, 2026, state that:

  • users who publish or disseminate outputs should verify authenticity and accuracy;
  • outputs may include information from third-party websites or external sources;
  • DeepSeek does not guarantee the validity, authenticity, legality, or security of external content;
  • outputs may contain errors or omissions and are for reference;
  • enabling Search may improve accuracy to some extent, but inaccuracy cannot be entirely avoided;
  • outputs used for decisions with legal or material effects on people require human review.

Those terms do not prove that a particular citation is right or wrong. They explain why a visible source marker should start verification rather than end it.

The broader risk is not unique to one product. The NIST Generative AI Profile treats confidently presented false, misleading, or unsupported content as a risk requiring testing, evaluation, verification, and validation. NIST does not provide a DeepSeek score and is not cited as an endorsement of this benchmark.

Privacy note before using DeepSeek Search

DeepSeek’s Privacy Policy, updated February 10, 2026, says the company integrates third-party APIs for search and shares input keywords to provide search services. It also describes collection of prompts, uploaded files, chat history, device and network information, and feature-use logs. The policy says the services are not designed or intended to process sensitive personal data and asks users not to provide it.

Use public, non-sensitive search questions unless your organization has approved a different workflow. Do not put confidential client names, unreleased products, private incidents, medical details, personal records, trade secrets, or unpublished research into a hosted search prompt merely because the answer may contain citations.

Our DeepSeek privacy policy guide explains the hosted-data boundary in more detail. For governance and high-risk decisions beyond citation quality, see the guide to whether DeepSeek is safe.

How to verify a DeepSeek citation in six steps

  1. Open the exact cited URL. Do not verify a source from its title, citation label, or search snippet alone.
  2. Confirm the source identity. Check the publisher, author or agency, document title, and final destination after redirects.
  3. Check the date and version. Look for an effective date, update date, named revision, historical year, correction, or superseding notice.
  4. Find support for the exact claim. Split compound sentences and verify every number, unit, date, qualifier, exception, and comparison.
  5. Prefer the original record. Replace an aggregator or derivative summary with the law, standard, dataset, official announcement, paper, or agency page when available.
  6. Record uncertainty. Remove or replace a citation that cannot be verified. Do not convert a paywall, bot block, or temporary failure into a confident support judgment.

For an academic reference, verify its DOI and bibliographic metadata through the publisher, Crossref, or an appropriate scholarly database, then read the source. For a current rule or statistic, confirm the effective date and population definition. For market and competitor work, use a source-checked market research workflow rather than copying a generated answer directly into a decision document.

Prompt Template for More Verifiable DeepSeek Search Citations

The following prompt cannot guarantee correct citations, but it makes the evidence boundary explicit:

Use Search and answer only with claims you can support from sources you opened. Prefer original or official sources. Place a citation immediately after every factual sentence. For each source, provide the title, publisher, date, and direct URL. If a source does not explicitly support a clause, remove that clause. If evidence conflicts or is insufficient, state the conflict or uncertainty. Do not invent a title, author, date, DOI, quotation, statistic, or URL.

For a multi-part task, add:

Separate each independently checkable claim. Preserve units, denominators, dates, versions, exceptions, and scope. After drafting, create a claim-to-source table and mark any claim that lacks full support.

Treat the table as an audit queue, not proof. A model can incorrectly approve its own citation, so the final check must open the source outside the answer.

When DeepSeek Search is not enough

DeepSeek Search may help discover and summarize public sources. It is not a substitute for:

  • a systematic or exhaustive literature search;
  • a qualified legal, medical, financial, tax, security, or standards review;
  • a licensed database needed to inspect controlling evidence;
  • document version control and approval;
  • privacy review for confidential or regulated material;
  • source permissions and access control;
  • independent human validation of a consequential claim.

If an organization needs answers grounded in an approved private corpus, a controlled retrieval system is a different problem. Our DeepSeek RAG knowledge base guide covers retrieval, permissions, source identifiers, evaluation, and auditability. It should not be confused with the hosted consumer Search feature tested here.

Students should follow institutional rules and cite evidence they verified, not a chatbot as a substitute for that evidence. The DeepSeek guide for students covers that broader workflow.

Protocol deviations

Three operator-side deviations were recorded across both protocols. In the July B02 baseline, a second identical prompt was accidentally submitted after the first response appeared absent in an accessibility snapshot. Only the untouched first completed response was frozen and scored. In the August expansion, the original US09 and IN11 captures ended before a complete answer was preserved. Each incomplete capture was excluded, and the identical prompt was rerun once in a fresh chat; the first completed rerun became the canonical scored answer.

BatchDeviationWhy it occurredEffect on scoring or interpretation
B02A second identical prompt was accidentally submittedThe first response was not visible in an apparently empty accessibility snapshotOnly the untouched first response was frozen; the duplicate was excluded and was never counted or scored
US09The first capture ended before the answer completedOperator-side evidence capture ended earlyThe incomplete capture was excluded; one identical-prompt rerun supplied the first complete answer and was scored
IN11The first capture ended before the answer completedOperator-side evidence capture ended earlyThe incomplete capture was excluded; one identical-prompt rerun supplied the first complete answer and was scored

The protocol allowed one clean rerun only when an operator-side capture failure prevented a complete answer from being preserved. A completed but weak, partially uncited, or inaccurate response remained the scored first response. No retry was allowed merely because an answer was poor.

In B04, screenshot capture timed out twice. That was an evidence-capture limitation, not a second scored-response deviation, because the live text, rendered HTML, and citation inventory were preserved.

Limitations of this citation audit

This was a controlled editorial benchmark, not a permanent product certification.

  • It used two dated test periods in a signed-in consumer web session; the account plan and identifier were not published.
  • Search results can vary by time, region, language, index state, account, and product update.
  • Only first completed responses were scored; another generation could differ.
  • The 48 English questions covered six global evidence domains plus U.S. and Indian official-source tasks, but they were not exhaustive.
  • The benchmark emphasized public sources with independently verifiable ground truth.
  • Paywalls, bot blocks, regional restrictions, and transient server errors can limit support adjudication.
  • Atomic-claim segmentation, support labels, and authority tiers require reviewer judgment.
  • For the July baseline, one primary reviewer completed the publication scoring. An independent checkpoint cross-checked all preserved baseline answers, required claims, citation occurrences, URL denominators, and the excluded B02 duplicate, but it did not complete a de novo page-level entailment rating. No V1 inter-rater agreement statistic is available.
  • That independent checkpoint covered the July baseline. The August country expansion received a separate source-by-source adjudication and a consistency QA pass, but not a second de novo blind rater; no V2 inter-rater agreement statistic is available.
  • The factual checkpoint initially differed from the primary ledger by one weighted point across five claims. Reconciliation retained 115.5/132 because the primary labels followed the frozen exact-qualifier and typical-use wording.
  • The 169-edge reconciliation uses source-specific claim-level decisions for B01-B03 and repairs edge multiplicity while preserving occurrence-level support labels for B04-B06. A fully de novo second source-entailment review of B04-B06 could change individual edge labels.
  • The independent checkpoint questioned the occurrence-level Tier A label assigned to a Washington State procurement PDF for a WCAG claim. Q17’s claim-level authority credit also had direct W3C support, so that flag did not change the question result.
  • The test does not establish behavior for DeepSeek’s API, open-weight models, third-party hosts, no-search mode, uploaded private files, other accounts, other regions, or future versions.
  • The U.S. and India labels identify the official-source subject matter. They do not prove the physical location of the browser, retrieval system, or model backend.
  • The July and August protocols have different chat units and therefore retain separate macro scores; no combined 48-question score is reported.
  • No independent country-subset PASS rule was preregistered; country score files label the subsets descriptively under inherited V1 gates while this article foregrounds question outcomes and exact metrics.
  • The V2 prompts did not request an exact question-ID heading, but the inherited frozen rubric included one format point for it. Both subsets lost the component consistently without an additional severity deduction.
  • The public release makes the benchmark rerunnable and the aggregate calculations reproducible, but it excludes private chat URLs, full generated answers, and raw interface evidence. Independent readers therefore cannot re-adjudicate every historical citation edge from the public package alone.
  • Citation performance does not measure every aspect of answer usefulness, bias, safety, completeness, or reasoning quality.

The baseline preferred sources were verified on July 26, 2026, and the country-expansion sources were verified on August 6, 2026. Historical-year and named-document questions remained anchored to the date or version in the prompt. Source drift found during scoring was recorded rather than silently changing the gold standard.

Final verdict: can DeepSeek cite sources accurately?

DeepSeek can produce accurate, well-cited answers, but neither dated protocol was reliable enough for unverified research or consequential decisions.

The July baseline reached 87.5% weighted factual accuracy but only 56.8182% weighted citation completeness. The U.S. expansion passed 4 of 12 questions with 89.2473% factual accuracy and 75.2688% completeness. The India expansion passed 0 of 12, with 68.6869% factual accuracy, 34.3434% completeness, and one critical reversal of current UIDAI guidance. The safest use is source discovery followed by independent, claim-by-claim verification.

Do not generalize the result beyond the tested consumer Search configuration and date. A DeepSeek citation should remain a verifiable pointer to evidence, not a substitute for reading the evidence.

For more practical workflows, browse the DeepSeek use-case library.

Frequently asked questions

Can DeepSeek cite sources accurately?

DeepSeek can display and sometimes correctly use sources when Search is enabled, but the 48 tested questions did not show dependable unverified citation performance. The July baseline passed 4/24 questions, the U.S. expansion passed 4/12, and the India expansion passed 0/12. The protocols retain separate scores because their chat units differ. Treat the citations as leads that require verification.

Does DeepSeek use live web search?

DeepSeek’s official app announcement lists Web Search as a feature, and its Terms refer to a user-enabled Search function. Availability and behavior can vary by product surface, account, region, and time. This audit tests only the documented consumer Chat configurations used on July 26 and August 6, 2026.

Does enabling Search make every DeepSeek answer accurate?

No. DeepSeek’s Terms say Search may improve accuracy to some extent, while also stating that inaccurate generated content cannot be entirely avoided. Weighted factual accuracy was 87.5% in the July baseline, 89.2473% in the U.S. expansion, and 68.6869% in the India expansion. Far fewer questions cleared every answer and citation gate.

What is the difference between a valid link and an accurate citation?

A valid link reaches the intended substantive page. An accurate citation must also support the exact attached claim, use an appropriate source and version, and be placed clearly. The July baseline found 79.1667% URL validity but 69.8225% support precision across 169 claim edges.

Does a 200 OK response prove that a citation is correct?

No. 200 OK means the request succeeded. It does not prove that the page is the intended document, that it is authoritative or current, or that its content supports DeepSeek’s claim.

Can DeepSeek invent citations or URLs?

Generative systems can produce plausible but incorrect references. The July baseline identified no fabricated URL, but it still found 8 invalid and 12 unverifiable citation occurrences. The country expansion also found no fabricated URL, while keeping invalid and unverifiable states separate. Do not label a citation fabricated solely because an automated request failed; redirects, bot blocks, paywalls, and temporary errors require separate review.

Are DeepSeek citations reliable for academic research?

Treat them as discovery leads, not final references. Verify the title, authors, publication, year, DOI, and source record through the publisher or a scholarly index, then read the work to confirm it supports the claim. Follow institutional and publisher rules for AI use and disclosure.

Does the DeepSeek API automatically provide web citations?

Do not infer API behavior from the consumer Chat Search interface. Official API documentation and the tested endpoint must establish any retrieval or citation feature. This audit does not score the API, open-weight models, or third-party DeepSeek hosts.

How do I verify a DeepSeek source?

Open the exact URL, confirm source identity, check the date and version, locate support for every adjacent clause, and prefer the original record when available. Remove or replace a source that cannot be verified, and record uncertainty instead of guessing.

What metrics should a citation audit report?

A transparent audit should report URL resolution, target identity, claim-level support, citation completeness, source authority, freshness, placement clarity, and answer correctness. Each rate needs a published denominator so readers can distinguish citation occurrences, citation-claim edges, weighted claims, questions, and batches.

How should redirects, paywalls, and blocked pages be scored?

Score them as separate states. A relevant redirect can pass after its destination is checked. A paywalled or bot-blocked source may be identifiable but not fully verifiable. A dead page, irrelevant redirect, or deceptive destination is a different failure. Automated inaccessibility alone does not prove fabrication.

Are primary sources always better than secondary sources?

Primary sources are usually preferable for official rules, specifications, statistics, court decisions, research findings, and company announcements. Strong secondary reporting can add context or challenge a primary claim. The right test is whether the source is authoritative, current, and appropriate for the exact statement.

Can one citation support several claims in the same sentence?

Yes, but every associated clause must be tested. The July baseline expanded 96 citation occurrences into 169 citation-claim edges because compound sentences often contained several atomic claims. One page may support all, some, or none of them.

How can I prompt DeepSeek to provide more verifiable sources?

Tell it to use Search, prefer primary sources, place a citation immediately after each factual sentence, provide direct URLs and dates, and state when evidence is insufficient. Instruct it not to invent titles, authors, quotations, DOIs, statistics, or links. Better prompting improves traceability but does not replace human verification.

Should I cite DeepSeek or the original source?

For factual claims, cite the original source you independently verified whenever possible. Cite or disclose DeepSeek separately when a style guide, institution, publisher, employer, or research protocol requires disclosure of AI assistance. Do not use the chatbot as a substitute for the evidence it points to.