# V2 Terra 02

- Object ID: `em:research-note:sha256:149a0523e912c1b8d35ebf5ef9ba9cba93bffde3416ee14b06ce4e0fc774b70e`
- Kind: `research-note`
- Repository path: [`research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-02.json`](https://github.com/yoheinakajima/epistemedia/blob/f92846570180dfa4511263f8ba98ecd18f7772c9/research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-02.json)
- Content digest: `819d320162a09909a73df00a4e1ebbdd42e67794cb02e3f7c379f601e10a0192`

**Also filed under:** [Research Program](https://epistemedia.org/topics/research-program/)

## Source content

{"question":"What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?","cutoff":"2026-08-22","answer":"Verified primary evidence exists. Four independent benchmark papers directly evaluate claim--URL support in deep-research-system outputs; LiveResearchBench is the only identified source here that explicitly separates non-resolving/inaccessible URLs, irrelevant URLs, and URLs that resolve but do not support the linked claim. The evidence is method- and snapshot-specific, not a general result about all agents.","results":[{"result_id":"r1_deepresearchbench_fact","proposition":"DeepResearch Bench's FACT evaluation found differing citation-support precision and supported-citation volume among four commercial deep-research agents.","reported_value":{"Grok Deeper Search":{"citation_accuracy_percent":83.59,"average_effective_supported_statement_url_pairs_per_task":8.15},"Perplexity Deep Research":{"citation_accuracy_percent":90.24,"average_effective_supported_statement_url_pairs_per_task":31.26},"Gemini-2.5-Pro Deep Research":{"citation_accuracy_percent":81.44,"average_effective_supported_statement_url_pairs_per_task":111.21},"OpenAI Deep Research":{"citation_accuracy_percent":77.96,"average_effective_supported_statement_url_pairs_per_task":40.79},"numerator_denominator":"Per-task supported unique statement--URL pairs / all unique statement--URL pairs, then macro-averaged across tasks; aggregate numerators and denominators are unknown."},"scope":{"model_agent":["Grok Deeper Search","Perplexity Deep Research","Gemini-2.5-Pro Deep Research","OpenAI Deep Research"],"dataset_population":"100 PhD-level tasks, 50 Chinese and 50 English, across 22 domains","tool_retrieval_path":"The FACT pipeline extracted and deduplicated statement--URL pairs, retrieved webpage text with Jina Reader API, and used Gemini-2.5-Flash for extraction and binary support judgment.","time":"Outputs collected in 2025: OpenAI 2025-04-01 to 2025-05-08; Gemini and Grok 2025-04-27 to 2025-04-29; Perplexity 2025-04-01 to 2025-04-29.","metric_scope":"Citation accuracy is macro-average support precision over unique statement--URL pairs; effective citations are supported pairs per task.","relevant_failure_class":["real source supporting a weaker claim","inaccessible source"]},"source_ids":["s1_deepresearchbench"],"exact_span_ids":["s1_fact_method","s1_fact_results","s1_fact_formula","s1_fact_judge_validation"],"interpretation":"This is direct evidence about whether cited webpages support extracted claims. It does not publish a separate invalid-URL/non-resolving-URL rate, so its reported non-support rate should not be interpreted as a clean decomposition of resolving versus inaccessible or irrelevant citations."},{"result_id":"r2_liveresearchbench_wide_info","proposition":"LiveResearchBench found non-zero citation errors for three leading systems on its Wide Info Search tasks, with unsupported claims the largest reported error class for each system.","reported_value":{"GPT-5":{"nonresolving_or_inaccessible_URL_errors_per_report":4.2,"irrelevant_URL_errors_per_report":1.7,"unsupported_claim_errors_per_report":13.3,"total_errors_per_report":19.2},"Grok-4 Deep Research":{"nonresolving_or_inaccessible_URL_errors_per_report":6.8,"irrelevant_URL_errors_per_report":6.8,"unsupported_claim_errors_per_report":33.4,"total_errors_per_report":47.0},"Open Deep Research":{"nonresolving_or_inaccessible_URL_errors_per_report":5.0,"irrelevant_URL_errors_per_report":5.2,"unsupported_claim_errors_per_report":19.7,"total_errors_per_report":29.9},"numerator_denominator":"Average errors per generated report; underlying counts and report denominator are unknown."},"scope":{"model_agent":["GPT-5","Grok-4 Deep Research","Open Deep Research"],"dataset_population":"The benchmark's Wide Info Search task category; exact task count used for Table 7 is unknown.","tool_retrieval_path":"An agentic judge grouped claims sharing a URL, checked whether the URL resolved, performed a coarse relevance check, then judged statement support from fetched content.","time":"Paper edition arXiv v5, 2026-04-18; individual system-output collection dates for this Table 7 subset are unknown.","metric_scope":"Citation errors, not an overall citation-accuracy percentage. E1 is inaccessible/non-resolving URL, E2 is irrelevant URL content, E3 is a resolving relevant URL that does not support the linked statement.","relevant_failure_class":["non-resolving URL","inaccessible source","irrelevant source","real source supporting a weaker claim"]},"source_ids":["s2_liveresearchbench"],"exact_span_ids":["s2_rubric_tree","s2_table7"],"interpretation":"This is the clearest identified direct evidence separating URL resolution/access, topical irrelevance, and claim-level non-support. It measures selected top performers and two task categories only."},{"result_id":"r3_liveresearchbench_market_analysis","proposition":"The same LiveResearchBench evaluation found higher average citation-error totals on Market Analysis than Wide Info Search for all three evaluated systems.","reported_value":{"GPT-5":{"nonresolving_or_inaccessible_URL_errors_per_report":11.1,"irrelevant_URL_errors_per_report":10.1,"unsupported_claim_errors_per_report":43.8,"total_errors_per_report":65.0},"Grok-4 Deep Research":{"nonresolving_or_inaccessible_URL_errors_per_report":6.3,"irrelevant_URL_errors_per_report":6.4,"unsupported_claim_errors_per_report":61.5,"total_errors_per_report":74.2},"Open Deep Research":{"nonresolving_or_inaccessible_URL_errors_per_report":11.9,"irrelevant_URL_errors_per_report":11.6,"unsupported_claim_errors_per_report":68.4,"total_errors_per_report":91.9},"comparison":"Table 7 reports lower total errors for the same systems on Wide Info Search: 19.2, 47.0, and 29.9, respectively.","numerator_denominator":"Average errors per generated report; underlying counts and report denominator are unknown."},"scope":{"model_agent":["GPT-5","Grok-4 Deep Research","Open Deep Research"],"dataset_population":"The benchmark's Market Analysis task category; exact task count used for Table 7 is unknown.","tool_retrieval_path":"Same rubric-tree agentic citation verifier as r2_liveresearchbench_wide_info.","time":"Paper edition arXiv v5, 2026-04-18.","metric_scope":"E1 inaccessible/non-resolving URL; E2 irrelevant URL; E3 unsupported linked claim.","relevant_failure_class":["non-resolving URL","inaccessible source","irrelevant source","real source supporting a weaker claim"]},"source_ids":["s2_liveresearchbench"],"exact_span_ids":["s2_table7"],"interpretation":"This is a second task-category result from the same underlying study, not independent corroboration of r2."},{"result_id":"r4_deeptrace","proposition":"DeepTRACE measured citation accuracy and unsupported-statement rates for public deep-research configurations, finding substantial variation rather than universal success or failure.","reported_value":{"GPT-5_DR":{"citation_accuracy_percent":79.1,"unsupported_relevant_statements_percent":12.5,"citation_thoroughness_percent":87.5},"YouChat_ARI":{"citation_accuracy_percent":39.33,"unsupported_relevant_statements_percent":62.85,"citation_thoroughness_percent":96.77},"YouChat_DR":{"citation_accuracy_percent":72.3,"unsupported_relevant_statements_percent":74.6,"citation_thoroughness_percent":83.5},"Perplexity_DR":{"citation_accuracy_percent":58.0,"unsupported_relevant_statements_percent":97.5,"citation_thoroughness_percent":9.1},"Copilot_Think_Deeper":{"citation_accuracy_percent":62.1,"unsupported_relevant_statements_percent":90.2,"citation_thoroughness_percent":13.2},"Gemini_DR":{"citation_accuracy_percent":50.3,"unsupported_relevant_statements_percent":53.6,"citation_thoroughness_percent":27.1},"numerator_denominator":"Citation accuracy = cited statement--source pairs judged supported / cited statement--source pairs. Unsupported statements = relevant statements unsupported by any listed source / relevant statements. Aggregate pair and statement counts are unknown."},"scope":{"model_agent":["GPT-5 Deep Research","YouChat ARI","YouChat Deep Research","Perplexity Deep Research","Copilot Think Deeper","Gemini Deep Research"],"dataset_population":"303 questions: 168 debate questions and 135 expert-contributed multi-search/hop questions; most metrics across 2,727 outputs (303 questions x 9 systems).","tool_retrieval_path":"Browser scripts extracted answer, citations, and public URLs; Jina Reader fetched source text. GPT-5 judged each statement--source factual-support matrix cell.","time":"Evaluation reported as of 2025-08-27.","metric_scope":"End-to-end statement--source support and citation placement; approximately 15% of URLs returned Jina Reader errors and were excluded from metrics requiring source full text.","relevant_failure_class":["inaccessible source","real source supporting a weaker claim","duplicate/shared source"]},"source_ids":["s3_deeptrace"],"exact_span_ids":["s3_definition","s3_table1","s3_corpus_and_retrieval","s3_judge_validation"],"interpretation":"This directly tests whether cited sources support statements, but does not separately count invalid/non-resolving URLs in its headline results. Its support decision was an LLM judgment with moderate reported correlation to a 100-item manual sample."},{"result_id":"r5_researcherbench","proposition":"ResearcherBench's factual-assessment pipeline found high cited-claim support precision but substantially lower citation coverage for several deep research systems on frontier-AI questions.","reported_value":{"OpenAI Deep Research":{"faithfulness_supported_cited_claims_over_cited_claims":0.84,"groundedness_cited_claims_over_all_claims":0.34},"Gemini Deep Research":{"faithfulness_supported_cited_claims_over_cited_claims":0.86,"groundedness_cited_claims_over_all_claims":0.59},"Grok3 DeepSearch":{"faithfulness_supported_cited_claims_over_cited_claims":0.69,"groundedness_cited_claims_over_all_claims":0.32},"Grok3 DeeperSearch":{"faithfulness_supported_cited_claims_over_cited_claims":0.80,"groundedness_cited_claims_over_all_claims":0.31},"Perplexity Deep Research":{"faithfulness_supported_cited_claims_over_cited_claims":0.85,"groundedness_cited_claims_over_all_claims":0.56},"numerator_denominator":"For each report, faithfulness = cited claims judged supported / cited claims; groundedness = cited claims / all extracted factual claims. Aggregate counts are unknown."},"scope":{"model_agent":["OpenAI Deep Research","Gemini Deep Research powered by Gemini-2.5-Pro","Grok3 DeepSearch","Grok3 DeeperSearch","Perplexity Deep Research"],"dataset_population":"65 frontier-AI research questions across 35 subjects, divided into technical-details, literature-review, and open-consulting questions.","tool_retrieval_path":"The pipeline extracted claim--context--URL triplets, used Jina Reader to extract URL text, then GPT-4.1 made a binary support decision.","time":"Evaluations conducted March--April 2025.","metric_scope":"Scientific/AI research reports; cited-claim support and citation coverage, not a separately reported URL-resolution rate.","relevant_failure_class":["real source supporting a weaker claim","inaccessible source"]},"source_ids":["s4_researcherbench"],"exact_span_ids":["s4_method","s4_results"],"interpretation":"This is direct claim-support evidence in a narrow frontier-AI research population. The apparent combination of high cited-claim faithfulness and low groundedness means it should not be read as evidence that most factual claims in the reports were supported."}],"sources":[{"source_id":"s1_deepresearchbench","url":"https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf","title":"DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents","authors_or_org":"Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao","date":"2025-06-13","identifier":"arXiv:2506.11763; DOI:10.48550/arXiv.2506.11763","edition":"arXiv v1, 31 pages","retrieval_status":"publicly retrieved","media_type":"research-paper PDF","license":"CC BY 4.0","exact_spans":[{"span_id":"s1_fact_method","locator":"p.4, Section 3.2, lines 148-155","quote":"\"Each unique Statement-URL pair undergoes a support evaluation.\"","supports":"FACT extracts, deduplicates, fetches webpage text, and judges binary claim support."},{"span_id":"s1_fact_results","locator":"p.5, Table 1, Deep Research Agent rows","quote":"\"Perplexity Deep Research ... 90.24 31.26\"","supports":"The Table 1 citation-accuracy and effective-citation values; the other three DRA rows are in the same table."},{"span_id":"s1_fact_formula","locator":"pp.17-18, Appendix E, equations 4-6","quote":"\"the proportion of 'support' statement-URL pairs for each individual task\"","supports":"Macro-average citation-accuracy definition and effective-citation calculation."},{"span_id":"s1_fact_judge_validation","locator":"p.17, Appendix C, lines 704-708","quote":"\"aligned with human 'support' determinations in 96% of cases\"","supports":"Reported 100-pair judge-validation result, including 92% agreement on not-support determinations."}]},{"source_id":"s2_liveresearchbench","url":"https://arxiv.org/pdf/2510.14240v5","title":"LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild","authors_or_org":"Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, Shafiq Joty","date":"2026-04-18","identifier":"arXiv:2510.14240v5; DOI:10.48550/arXiv.2510.14240","edition":"arXiv v5; accepted to ICLR 2026","retrieval_status":"publicly retrieved","media_type":"research-paper PDF and official public code repository","license":"CC BY 4.0 for arXiv paper; repository code Apache-2.0","exact_spans":[{"span_id":"s2_rubric_tree","locator":"p.33, Appendix E","quote":"\"E1: URL is inaccessible or does not resolve; E2: URL content is irrelevant ... E3: URL content does not support the specific statements.\"","supports":"Exact failure-class definitions and the resolving/relevance/support verification sequence."},{"span_id":"s2_table7","locator":"p.33, Table 7","quote":"\"GPT-5 4.2 1.7 13.3 19.2\"","supports":"Wide Info Search error counts; Table 7 also supplies the Grok-4, Open Deep Research, and Market Analysis rows."},{"span_id":"s2_validation","locator":"p.23, Appendix C, Citation Accuracy","quote":"\"human evaluators agree with either Gemini or GPT-5 on 87.1%\"","supports":"Reported human agreement on 200 sampled claim--URL support judgments."}]},{"source_id":"s3_deeptrace","url":"https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf","title":"DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence","authors_or_org":"Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, Chien-Sheng Wu","date":"2026 (exact conference-publication day unknown)","identifier":"ICLR 2026 paper ID ad08767706825033b99122332293033d","edition":"Published ICLR 2026 conference paper, 22 pages","retrieval_status":"publicly retrieved","media_type":"conference-paper PDF","license":"unknown","exact_spans":[{"span_id":"s3_definition","locator":"p.6, Section 3.1.4, equation 7","quote":"\"fraction of statement citations that accurately reflect that a source's content supports the statement\"","supports":"Citation-accuracy metric definition."},{"span_id":"s3_table1","locator":"p.8, Table 1","quote":"\"%Citation Accuracy 79.1 ... 39.33 ... 72.3 ... 58.0 ... 62.1 ... 50.3\"","supports":"Reported DR-system citation-accuracy values and adjacent unsupported-statement/thoroughness rows."},{"span_id":"s3_corpus_and_retrieval","locator":"pp.4-7, Sections 3.1.1 and 3.2","quote":"\"The dataset comprises 303 questions\"","supports":"Population, public-UI collection, Jina Reader path, and stated source-extraction exclusions."},{"span_id":"s3_judge_validation","locator":"p.14, Table 3","quote":"\"Factual support (statement–source) 0.62\"","supports":"Reported Pearson correlation between LLM factual-support labels and human annotations, N=100 per task."}]},{"source_id":"s4_researcherbench","url":"https://arxiv.org/html/2507.16280","title":"ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry","authors_or_org":"Tianze Xu, Pengrui Lu, Lyumanshan Ye, Xiangkun Hu, Pengfei Liu","date":"2025-07-22","identifier":"arXiv:2507.16280v1; DOI:10.48550/arXiv.2507.16280","edition":"arXiv v1","retrieval_status":"publicly retrieved","media_type":"research-paper HTML","license":"arXiv.org perpetual non-exclusive license","exact_spans":[{"span_id":"s4_method","locator":"Section 4.2.1, Citation Support Verification and Score Computation","quote":"\"whether the extracted content supports the corresponding claim\"","supports":"Claim--URL--context extraction, Jina Reader retrieval, binary support decision, and faithfulness/groundedness definitions."},{"span_id":"s4_results","locator":"Section 5.2, Table 2","quote":"\"OpenAI Deep Research | 0.7032 | 0.84 | 0.34\"","supports":"Coverage, faithfulness, and groundedness values; the other systems appear in the same table."}]}],"counterevidence":[{"claim":"No study above establishes that deep-research-agent citations are generally unreliable or generally reliable.","basis":"The point estimates differ materially across benchmark populations, product snapshots, output formats, source-fetch paths, extraction rules, and LLM judges. FACT reports 77.96%-90.24% macro citation precision for four DRAs, whereas DeepTRACE reports 50.3%-79.1% citation accuracy for overlapping product families under a different corpus and metric."},{"claim":"A high cited-claim support rate is not evidence that an entire report is supported.","basis":"ResearcherBench reports faithfulness of 0.84-0.86 for OpenAI/Gemini Deep Research but groundedness of 0.34/0.59; FACT separately reports effective supported citations per task."},{"claim":"The DeepTRACE Gemini DR citation-accuracy value has an internal artifact inconsistency.","basis":"Table 1 reports 50.3%, while the surrounding prose states 40.3%. This response retains the table's exact number and does not resolve the discrepancy."}],"limitations":["All four evaluations rely materially on an LLM judge after webpage-text extraction; their reported agreement checks do not establish error-free support labels.","Jina Reader extraction can fail or omit page content. DeepTRACE states roughly 15% of URLs returned an error and excludes them from full-text-dependent calculations; this can bias support metrics and does not make inaccessible citations harmless.","DeepResearch Bench and ResearcherBench do not provide a separate published numerator/rate for non-resolving URLs, inaccessible sources, irrelevant URLs, duplicates, or weaker-than-claimed support.","LiveResearchBench Table 7 reports error counts per report, not claim-level denominators or rates, and examines only three selected systems on two task categories.","System labels are product snapshots, not stable model weights or reproducible runtime configurations. DeepResearch Bench explicitly notes commercial iteration opacity.","The cited papers are separate empirical studies; their results must not be treated as independent measurements when merely slicing the same study by task category, model profile, or metric."],"unresolved":["The exact Table 7 task/report denominator and raw error counts in LiveResearchBench are unknown from the published table.","Whether individual Jina Reader failures were invalid URLs, paywalls, robots restrictions, transient failures, or parser failures is generally unknown.","For the listed papers, public raw claim--URL judgment ledgers sufficient to independently recompute every published value were not verified here.","No verified source in this set tests every named failure class simultaneously for every agent, at every time, with human gold labels.","The precise conference-publication day and license for the DeepTRACE conference PDF are unknown."],"search_notes":{"method":"Searched public, credential-free web and official paper/project pages for deep-research citation accuracy, claim--URL support, URL resolution, inaccessible URL, irrelevant URL, and unsupported-claim evaluations. Excluded commentary, search snippets, unverified repositories, and benchmarks that measured only retrieval recall or report quality without claim--citation support testing.","inclusion_rule":"Included only primary research papers that publicly describe an empirical claim--URL support method and report agent/system results by the cutoff.","non_independence_note":"LiveResearchBench Wide Info Search and Market Analysis are two result rows from one paper and one verifier; they are reported separately for scope clarity but are not independent studies."}}
