# V2 Terra 03

- Object ID: `em:research-note:sha256:be7444c886f2485242f30836e3a897b5bbcd59b7f66507c4fd92f78a3e06f898`
- Kind: `research-note`
- Repository path: [`research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-03.json`](https://github.com/yoheinakajima/epistemedia/blob/f92846570180dfa4511263f8ba98ecd18f7772c9/research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-03.json)
- Content digest: `e211dfe49066bb9f4ff9e59b91173d3b0f45b7b1e8ca5135c7435f3882d7e264`

**Also filed under:** [Research Program](https://epistemedia.org/topics/research-program/)

## Source content

{"question":"What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?","cutoff":"2026-08-22","answer":"Verified primary evidence exists, but it is limited, method-dependent, and does not establish a universal rate for deep-research agents. The most direct study found high URL accessibility and topical relevance but substantially lower claim-to-source factual support across 14 web-search research agents. Earlier benchmarks independently measured statement-URL support or semantic consistency, and a small biomedical study manually checked both bibliographic identifiability and claim-citation concordance. These are separate experiments; the 2026 source-attribution study reuses DeepResearch Bench tasks for part of its query set, so it is not independent task evidence from that benchmark.","results":[{"result_id":"R1","proposition":"In a 14-model web-search research-agent experiment, working URLs and topical relevance were much higher than factual claim support.","reported_value":{"numerator":"unknown","denominator":"unknown","rate_or_comparison":{"frontier_link_works":"94.1%-100.0% by model","frontier_relevant_content":"80.7%-95.7% by model","frontier_fact_check":"38.9%-76.8% by model","full_table_fact_check_range":"24.4%-76.8% across 14 models"},"quantitative_basis":"Table 1 reports percentages; pair-level counts are not reported in the paper."},"scope":{"model_agent_dataset_tool_population_time_metric":"14 LLM web-search agents: GPT-5.2, GPT-5.4, GPT-5 Mini, Codex; Claude Sonnet 4.5/4.6, Opus 4.5/4.6, Haiku 4.5; Gemini 3.1 Pro and 3 Flash; OSS-120B, Llama 4 Maverick, Pixtral Large. 130 research queries drawn from DeepResearch Bench and BrowseComp. Each agent generated Markdown with inline citations under a citation/minimum-search-depth system prompt. Each citation-claim pair was parsed with a Markdown AST; URLs were fetched; Link Works was binary accessibility; Relevant Content and Fact Check were binary LLM-judge measures. Models and source availability were evaluated in 2026; exact collection dates are unknown."},"source_ids":["S1"],"exact_span_ids":["S1-A","S1-B","S1-C"],"interpretation":"Verified source fact: the reported factual-support metric was lower than the reported URL and topical metrics in this experiment. Interpretation: a resolvable, relevant citation did not by itself establish support for the attached claim in this tested setting; this is not a rate for all agents or all citations."},{"result_id":"R2","proposition":"Increasing permitted search depth reduced the reported factual-support score in a controlled two-model ablation while link accessibility and relevance stayed high.","reported_value":{"numerator":"unknown","denominator":"unknown","rate_or_comparison":{"GPT-5.4_fact_check":"78.6% at 2 calls versus 16.7% at 150 calls; decline 61.9 percentage points","Claude_Opus_4.6_fact_check":"80.0% at 2 calls versus 57.9% at 150 calls; decline 22.1 percentage points","reported_average_drop":"approximately 42% from 2 to 150 calls across the two models","link_and_relevant":"reported above 92% at all tested depths"},"quantitative_basis":"Tables 2-3; the paper does not publish the underlying number of citation-claim pairs at each depth."},"scope":{"model_agent_dataset_tool_population_time_metric":"GPT-5.4 and Claude Opus 4.6, at maximum tool-call caps of 2, 10, 30, 50, 70, 100, and 150, using the S1 source-attribution pipeline. Fact Check is source-content support/consistency judged after source retrieval; Link Works and Relevant Content are distinct metrics."},"source_ids":["S1"],"exact_span_ids":["S1-D","S1-E"],"interpretation":"Verified source fact: factual-support scores declined in this particular controlled ablation. Interpretation: this supports a hypothesis of synthesis overload for these two configurations, not a general causal claim that more research always worsens citations."},{"result_id":"R3","proposition":"DeepResearch Bench reported statement-URL support precision for four commercial deep-research agents on 100 expert-designed tasks.","reported_value":{"numerator":"unknown","denominator":"unknown","rate_or_comparison":{"Perplexity_Deep_Research_citation_accuracy":"90.24%; 31.26 effective supported statement-URL pairs per task","Gemini_2_5_Pro_Deep_Research_citation_accuracy":"81.44%; 111.21 effective supported statement-URL pairs per task","OpenAI_Deep_Research_citation_accuracy":"77.96%; 40.79 effective supported statement-URL pairs per task","Grok_Deeper_Search_citation_accuracy":"83.59%; 8.15 effective supported statement-URL pairs per task"},"quantitative_basis":"Citation Accuracy is the mean, across tasks, of support-judged unique statement-URL pairs divided by all unique statement-URL pairs; a task with none scores zero. Effective Citations is all support-judged pairs divided by 100 tasks. Raw pair counts are not reported."},"scope":{"model_agent_dataset_tool_population_time_metric":"100 PhD-level tasks, 50 Chinese and 50 English, across 22 domains. Evaluated agents: Gemini-2.5-Pro Deep Research, OpenAI Deep Research, Grok Deeper Search, and Perplexity Deep Research. Data collection: OpenAI 2025-04-01 to 2025-05-08; Gemini 2025-04-27 to 2025-04-29; Perplexity 2025-04-01 to 2025-04-29; Grok 2025-04-27 to 2025-04-29. Pages were retrieved through Jina Reader; Gemini-2.5-Flash extracted statement-URL pairs and made binary support judgments."},"source_ids":["S2","S5"],"exact_span_ids":["S2-A","S2-B","S2-C","S2-D","S5-A","S5-B"],"interpretation":"Verified source fact: this edition reports the stated support-precision and supported-pair-throughput estimates. Interpretation: these figures measure the benchmark's judge-defined support condition, not URL reachability, and the current official repository says its evaluator changed after publication; do not substitute later leaderboard scores for this v1 result."},{"result_id":"R4","proposition":"ReportBench reported semantic consistency between cited statements and retrieved cited webpages for two commercial deep-research products on academic-survey tasks.","reported_value":{"numerator":"unknown","denominator":"unknown","rate_or_comparison":{"OpenAI_Deep_Research_match_rate":"78.87%; 88.2 cited statements per report; 9.89 references per report","Gemini_Deep_Research_match_rate":"72.94%; 96.2 cited statements per report; 32.42 references per report","comparison":"OpenAI Deep Research exceeded Gemini Deep Research by 5.93 percentage points on reported cited-statement match rate."},"quantitative_basis":"Table 1 reports average statements/references and match-rate percentages, not raw verified-statement totals."},"scope":{"model_agent_dataset_tool_population_time_metric":"100 prompts reverse-engineered from peer-reviewed arXiv survey papers, selected from 678 retained papers and concentrated in STEM. Reports were manually collected from OpenAI and Gemini web interfaces during 2025-07-14 to 2025-07-25. The paper says standard OpenAI Deep Research was powered by o3; Gemini had Gemini 2.5 Pro and Deep Research toggles enabled. GPT-4o extracted statements, found support passages from web-scraped cited pages, and performed semantic consistency verification."},"source_ids":["S3"],"exact_span_ids":["S3-A","S3-B","S3-C","S3-D"],"interpretation":"Verified source fact: the paper reports the two match rates under its semantic-consistency pipeline. Interpretation: it is relevant evidence of claim-source alignment, but it does not separately report a non-resolving-URL rate and depends on automated extraction, retrieval, passage selection, and GPT-4o judgment."},{"result_id":"R5","proposition":"A preliminary biomedical evaluation checked bibliographic identifiability and citation-bearing-sentence concordance in five research systems and found both fabricated references and source-claim inaccuracies.","reported_value":{"numerator":"unknown","denominator":"unknown","rate_or_comparison":{"ChatGPT_Deep_Research_reference_identifiable":"95.7%","ChatGPT_Deep_Research_references_entirely_correct":"69.6%","ChatGPT_Deep_Research_one_or_two_minor_errors_but_identifiable":"21.7%","ChatGPT_Deep_Research_multiple_minor_errors_but_identifiable":"4.3%","fake_authors_or_titles":"Claude 95.8% +/- 7.2%; Gemini 47.6% +/- 7.8%; Perplexity 50.1% +/- 28.0%","citation_bearing_sentence_inaccuracy":"ChatGPT Deep Research online-search mode 50.0%; uploaded-paper mode 40.0%, as reported in Figure 1/search-indexed full text; raw sentence denominator unknown"},"quantitative_basis":"Three independent runs, 1,500-word limit, APA-style references; individual reference and sentence totals are not reported in the article text located."},"scope":{"model_agent_dataset_tool_population_time_metric":"Five systems with web access: ChatGPT Deep Research (o3-mini-high), Claude Opus 4 Research, Gemini 2.5 Pro Preview Deep Research, Le Chat/Mistral premier Think, and Perplexity.AI Best Deep Research. Task: survey baseline performance, enhancement strategies, and ethical considerations of ChatGPT in dermatological image analysis. Reference lists were checked for accuracy and citation-bearing sentences for claim-citation concordance; three independent runs. The article also compared online-search and ten-uploaded-paper modes for ChatGPT and Le Chat."},"source_ids":["S4","S6"],"exact_span_ids":["S4-A","S4-B","S4-C","S6-A"],"interpretation":"Verified source fact: the primary article reports the listed metadata-identifiability and reference-error results, and its figure reports sentence-level inaccuracy categories. Interpretation: this is direct but preliminary, single-topic evidence; bibliographic identifiability is not equivalent to a live URL resolving, and sentence errors can have more than one category."}],"sources":[{"source_id":"S1","url":"https://arxiv.org/html/2605.06635","title":"Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents","authors_or_org":"Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah, Corey Feld; Commercial Technology and Innovation Office, PricewaterhouseCoopers U.S.","date":"2026-05-07","identifier":"arXiv:2605.06635v1","edition":"v1 examined, arXiv HTML, dated 2026-05-07","retrieval_status":"public full text retrieved","media_type":"research preprint / HTML","license":"CC BY 4.0","exact_spans":[{"span_id":"S1-A","locator":"Abstract, lines 45-49","quote":"Citations are evaluated along three dimensions. (1) Link Works verifies URL accessibility, (2) Relevant Content measures topical alignment, and (3) Fact Check validates factual accuracy against source content.","supports":"metric definitions"},{"span_id":"S1-B","locator":"Table 1, lines 163-180","quote":"Model | Success | Link Works | Relevant | Fact Check","supports":"the model-specific percentages and full-range comparison in R1"},{"span_id":"S1-C","locator":"Methods 3.4, lines 153-156","quote":"We evaluate 14 LLMs ... on 130 research queries drawn from DeepResearch Bench and BrowseComp.","supports":"scope and shared-upstream-task warning"},{"span_id":"S1-D","locator":"Tables 2-3, lines 191-210","quote":"2 | 100.0% | 100.0% | 78.6% ... 150 | 99.2% | 99.2% | 16.7%","supports":"GPT-5.4 depth ablation"},{"span_id":"S1-E","locator":"Evaluation 4.3, lines 211-215","quote":"Fact Check accuracy drops approximately 42% on average from minimal (2 calls) to maximal search depth.","supports":"reported average depth comparison"},{"span_id":"S1-F","locator":"Limitations, lines 219-224","quote":"the LLM-as-a-judge approach ... may retain biases inherent to the judge model","supports":"judge and temporal-stability limitations"}]},{"source_id":"S2","url":"https://arxiv.org/html/2506.11763","title":"DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents","authors_or_org":"Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao","date":"2025-06-13","identifier":"arXiv:2506.11763v1","edition":"v1 examined, arXiv HTML, dated 2025-06-13","retrieval_status":"public full text retrieved","media_type":"research preprint / HTML","license":"CC BY 4.0","exact_spans":[{"span_id":"S2-A","locator":"Evaluation methodology, lines 144-149","quote":"Each unique Statement-URL pair undergoes a support evaluation.","supports":"support metric methodology"},{"span_id":"S2-B","locator":"Table 1, lines 177-182","quote":"Perplexity Deep Research | ... | 90.24 | 31.26 ... Gemini-2.5-Pro Deep Research | ... | 81.44 | 111.21 ... OpenAI Deep Research | ... | 77.96 | 40.79","supports":"R3 figures"},{"span_id":"S2-C","locator":"Appendix E, lines 404-423","quote":"Citation Accuracy ... [is] the proportion of 'support' statement-URL pairs ... averaging these per-task accuracies across all tasks.","supports":"quantitative basis"},{"span_id":"S2-D","locator":"Appendix A, lines 339-348","quote":"Each of our 100 tasks was developed by a verified PhD-level expert ... [but] inevitably constrains the dataset size.","supports":"scale and generalizability limitations"}]},{"source_id":"S3","url":"https://arxiv.org/html/2508.15804","title":"ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks","authors_or_org":"Minghao Li, Ying Zeng, Zhihao Cheng, Cong Ma, Kai Jia; ByteDance BandAI","date":"2025-08-14","identifier":"arXiv:2508.15804v1","edition":"v1 examined, arXiv HTML, dated 2025-08-14","retrieval_status":"public full text retrieved","media_type":"research preprint / HTML","license":"CC BY 4.0","exact_spans":[{"span_id":"S3-A","locator":"Cited-statements pipeline, lines 121-124","quote":"we retrieve the full content of each cited webpage ... [and] perform consistency verification by comparing the statement with the retrieved content","supports":"claim-source alignment method"},{"span_id":"S3-B","locator":"Settings and metrics, lines 132-139","quote":"For statement extraction, supporting source extraction, and semantic consistency verification, we adopt gpt-4o.","supports":"judge/tool scope"},{"span_id":"S3-C","locator":"Table 1, lines 140-150","quote":"OpenAI Deep Research | 0.385 | 0.033 | 9.89 | 78.87% | 88.2 ... Gemini Deep Research | 0.145 | 0.036 | 32.42 | 72.94% | 96.2","supports":"R4 quantitative figures"},{"span_id":"S3-D","locator":"Limitations, lines 259-261","quote":"The benchmark primarily draws from peer-reviewed survey papers on arXiv, most of which are concentrated in STEM fields.","supports":"domain limitation"}]},{"source_id":"S4","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/","title":"Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype","authors_or_org":"Lauren E. Keplinger, Luke K. Frashure, Sabrina A. Duran, Gangqing Hu","date":"2025-09-04","identifier":"DOI:10.1111/jdv.70035; PMID:40904191; PMCID:PMC13109748","edition":"electronic publication 2025-09-04; issue 2026-05, Journal of the European Academy of Dermatology and Venereology 40(5):e300-e302","retrieval_status":"public primary full text search-indexed; direct re-open during this trace returned a PMC automated-browser challenge","media_type":"peer-reviewed letter / preliminary empirical evaluation","license":"CC BY-NC 4.0","exact_spans":[{"span_id":"S4-A","locator":"Main text, paragraph beginning 'To address the gap' (search-indexed primary full text)","quote":"We imposed a 1500-word limit, APA-style references and conducted three independent runs.","supports":"study scope"},{"span_id":"S4-B","locator":"Main text, following paragraph (search-indexed primary full text)","quote":"ChatGPT Deep Research-generated references were mostly identifiable by their metadata (95.7%): 69.6% were entirely correct","supports":"identifiability and correctness figures"},{"span_id":"S4-C","locator":"Same paragraph (search-indexed primary full text)","quote":"fake authors or titles occurred frequently in references from Claude (95.8 +/- 7.2%), Gemini (47.6 +/- 7.8%) and Perplexity.AI (50.1 +/- 28.0%).","supports":"fabrication-class figures"}]},{"source_id":"S5","url":"https://github.com/Ayanami0730/deep_research_bench","title":"DeepResearch Bench official repository","authors_or_org":"Ayanami0730 / DeepResearch Bench project","date":"2026-05-11 update visible in repository","identifier":"GitHub repository Ayanami0730/deep_research_bench","edition":"repository main branch inspected on 2026-08-22; not the frozen paper-v1 evaluator","retrieval_status":"public repository retrieved","media_type":"official project artifact","license":"Apache-2.0","exact_spans":[{"span_id":"S5-A","locator":"README News, lines 176-188","quote":"Official Evaluator Switched to GPT-5.5 ... Legacy code: the previous Gemini-2.5-Pro / Gemini-2.5-Flash evaluation code is preserved","supports":"edition drift warning"},{"span_id":"S5-B","locator":"README FACT, lines 240-248","quote":"Support Verification: Uses web scraping and LLM judgment to verify whether cited sources actually support the claims","supports":"official description of FACT's intended function"}]},{"source_id":"S6","url":"https://pubmed.ncbi.nlm.nih.gov/40904191/","title":"PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype","authors_or_org":"U.S. National Library of Medicine record for Keplinger et al.","date":"2025-09-04 electronic publication; 2026-05 issue","identifier":"PMID:40904191","edition":"PubMed figure-caption view inspected","retrieval_status":"public record retrieved","media_type":"authoritative bibliographic/index artifact","license":"unknown","exact_spans":[{"span_id":"S6-A","locator":"Figure 1 caption","quote":"Six inaccuracy types were identified within citation-bearing sentences ... Sentences may bear multiple types of inaccuracy, simultaneously.","supports":"failure-class and non-exclusive-category caution"}]}],"counterevidence":[{"source_ids":["S1"],"verified_fact":"The strongest score in S1's Fact Check column was 76.8% for Claude Opus 4.5, while its Link Works and Relevant scores were 98.7% and 95.7%; the study does not report universal failure."},{"source_ids":["S2"],"verified_fact":"S2 reported 77.96%-90.24% citation accuracy for the four tested commercial deep-research agents under its FACT support definition."},{"source_ids":["S3"],"verified_fact":"S3 reported 72.94% and 78.87% cited-statement semantic consistency for Gemini and OpenAI Deep Research, respectively, in its academic-survey setting."},{"source_ids":["S4"],"verified_fact":"S4 reported that 95.7% of ChatGPT Deep Research references were identifiable by metadata, although identifiability did not establish all cited sentences were accurately supported."}],"limitations":["None of the reported rates is a general deep-research-agent rate; model, product runtime, prompts, tools, date, content type, retrieval permissions, and evaluator differ.","S1 is the most direct evidence that separately evaluates resolving URLs and claim support, but it uses LLM judges for relevance and Fact Check, calibrated only through manual review of 50-100 judgments; its agents are configured research agents, not necessarily consumer product modes.","S1's 130 queries are drawn from both DeepResearch Bench and BrowseComp. Therefore S1 and S2 share upstream tasks/data and should not be counted as independent task datasets, even though their model sets, runs, and support pipelines differ.","S2's reported paper results use Jina Reader and Gemini-2.5-Flash for extraction/support judgment; its official repository later switched evaluators. Paper-v1 values must not be compared numerically to later leaderboard values without a frozen edition and re-run.","S3's cited-statement match metric requires web scraping, LLM support-passage selection, and GPT-4o semantic verification. It reports no direct broken-link/non-resolving-URL rate. Its gold references are from survey-paper bibliographies, which can penalize valid sources absent from the reference set, although the claim-support metric is separate.","S4 is a short preliminary, narrow-domain study with three runs, a fixed 1,500-word/APA task, incompletely reported raw denominators, and non-exclusive sentence error categories. 'Identifiable by metadata' is not the same as a user-clickable URL resolving at evaluation time.","S4's primary full text is public but a direct automated re-open encountered a challenge in this trace; the quoted text is from the publicly indexed primary full text and PubMed metadata/figure caption. Its reported numbers should be independently rechecked against the archived article/supplement before a formal pooled analysis.","The available evidence primarily concerns online/web citations and does not establish results for internal RAG, scholarly reference managers, paywalled corpora, non-English web retrieval, or high-stakes professional use."],"unresolved":["Exact raw numerator and denominator counts for support/failure rates are unavailable in the located paper text for S1-S4.","'Resolve' has no single shared definition: S1 treats accessible extracted content as Link Works and includes HTTP failures, blocked access, timeouts, paywalls, and removed URLs; S4 measures bibliographic identifiability; S2 and S3 retrieve pages but do not publish an equivalent accessibility percentage.","'Support' differs materially: S1 Fact Check includes source-content consistency; S2 uses a binary support judge on deduplicated statement-URL pairs; S3 uses semantic consistency after a selected support passage; S4 uses author evaluation and six non-exclusive inaccuracy classes.","Whether sentence-end citations apply to one or multiple preceding claims is a live attribution-design issue. S1 explicitly uses backward attribution, so its pair counts and support scores depend on that rule.","No located source provides a human-adjudicated, longitudinal, cross-product benchmark simultaneously reporting URL resolution, access/paywall status, relevance, exact entailment/claim support, duplicate/shared-source dependence, and public raw outputs across agents."],"search_notes":{"included":"Only primary papers, their official repositories, and NLM/PMC indexing artifacts were used for empirical claims. Commentary, news, Reddit, and generated summaries were excluded as evidence.","excluded_or_not_counted":"BrowseComp and general attribution/fact-check benchmarks were not counted because they do not themselves measure citations emitted by deep-research agents. A 2025 perspective on oculomics is useful qualitative corroboration but was not counted as a quantitative result because it reports no exact citation-support denominator/rate. A Bratislava Medical Journal comparison was not counted because its own text says some values are indirect estimates rather than direct standardized tests.","search_paths":"Searched public arXiv, primary journal/PMC/PubMed records, and official GitHub project repositories for deep-research citation accuracy, citation consistency, source attribution, resolving links, and claim-support evaluations.","freshness":"All sources and examined editions predate the stated 2026-08-22 cutoff; web citations remain temporally unstable after the evaluation dates."}}
