Repository object · research-note

V2 Terra 01

Accepted research note in the public catalog.

Source path
research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-01.json
Media type
application/json
Object ID
em:research-note:sha256:01a8b250fdf427900cdd606698b6704c5c3bd28487e9e718113d365453c87c01
Content digest
e90c36b6f172fbe291d75ce1ea9451b1160515a2b7b4e0ead66aa1a76cc8b735

Source content

{"question":"What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?","cutoff":"2026-08-22","answer":"Verified primary studies measure the two properties separately and, in several cases, together. Claim-to-source support was measured on 100-task DeepResearch Bench, 100-task ReportBench, 65-question ResearcherBench, and 130-query Cited but Not Verified. URL resolution was separately measured at scale in Detecting and Correcting Reference Hallucinations. The results are not directly interchangeable: datasets, output vintages, citation parsing, source-access rules, judges, and denominators differ. They show both substantial supported-citation rates in particular tested conditions and material failure rates; none establishes a rate for all deep-research agents or current production versions.","results":[{"result_id":"r1_deepresearchbench_fact","proposition":"DeepResearch Bench FACT measured whether extracted, deduplicated statement-URL pairs had webpage evidence sufficient to support the statement.","reported_value":{"numerator":"unknown aggregate supported-pair count; each task numerator was supported unique statement-URL pairs","denominator":"100 tasks; each task denominator was its unique extracted statement-URL pairs","rate":"per-task support proportions averaged over tasks: Perplexity Deep Research 90.24%, Gemini-2.5-Pro Deep Research 81.44%, OpenAI Deep Research 77.96%, Grok Deeper Search 83.59%","comparison":"average effective supported statement-URL pairs per task: Gemini 111.21, OpenAI 40.79, Perplexity 31.26, Grok 8.15"},"scope":{"models_or_agents":["Perplexity Deep Research","Gemini-2.5-Pro Deep Research","OpenAI Deep Research","Grok Deeper Search"],"dataset_or_population":"100 expert-created PhD-level research tasks across 22 fields, 50 Chinese and 50 English","tool_and_retrieval_path":"Jina Reader retrieved cited webpage text; Gemini-2.5-Flash extracted/deduplicated statement-URL pairs and made binary support judgments","time":"commercial-agent outputs collected in 2025: OpenAI 2025-04-01 to 2025-05-08; Gemini 2025-04-27 to 2025-04-29; Perplexity 2025-04-01 to 2025-04-29; Grok 2025-04-27 to 2025-04-29","metric_scope":"citation accuracy is the mean of task-level supported-pair proportions, with a zero for a task having no citable statements; it is not URL liveness or all-claim coverage"},"source_ids":["s1_deepresearchbench"],"exact_span_ids":["s1_method","s1_table","s1_dates"],"interpretation":"Verified study result: support, not merely topical relevance, was evaluated for extracted pairs. The rates are benchmark-and-judge-specific rather than general properties of the products."},{"result_id":"r2_reportbench_cited_statement_match","proposition":"ReportBench measured semantic consistency of cited statements against retrieved content from their cited webpages.","reported_value":{"numerator":"unknown supported cited-statement count","denominator":"100 ReportBench reports/tasks; per-report cited-statement denominator not reported in the table","rate":"OpenAI Deep Research cited-statement Match Rate 78.87%; Gemini Deep Research 72.94%","comparison":"average cited statements per report: OpenAI 88.2 versus Gemini 96.2; reference precision 0.385 versus 0.145; reference recall 0.033 versus 0.036"},"scope":{"models_or_agents":["OpenAI Deep Research, standard version powered by o3","Gemini Deep Research with Gemini 2.5 Pro and Deep Research toggles enabled"],"dataset_or_population":"100 reverse-prompted academic survey tasks sampled from 678 formally published arXiv survey papers, across 11 categories","tool_and_retrieval_path":"gpt-4o extracted cited statements, scraped each cited webpage, located a semantically relevant passage, and performed consistency verification","time":"outputs manually collected from web interfaces 2025-07-14 to 2025-07-25","metric_scope":"Match Rate is the proportion of cited statements semantically consistent with cited sources; it does not establish HTTP/URL resolution rate, and it is distinct from reference overlap and non-cited-statement factual accuracy"},"source_ids":["s2_reportbench"],"exact_span_ids":["s2_method","s2_metric_table","s2_collection"],"interpretation":"Verified study result: in this particular academic-survey benchmark and automated pipeline, both tested agents had nonzero cited-statement mismatch rates; the result does not establish that one agent is generally better."},{"result_id":"r3_researcherbench_faithfulness_and_groundedness","proposition":"ResearcherBench measured whether a cited claim is supported by the textual content extracted from its citation URL, and separately measured citation coverage of all factual claims.","reported_value":{"numerator":"unknown aggregate supported cited-claim count; defined per question as Ns,k","denominator":"unknown aggregate cited-claim count; defined per question as Nc,k; benchmark contains 65 questions","rate":"faithfulness (supported cited claims / cited claims): OpenAI Deep Research 0.84, Gemini Deep Research 0.86, Grok3 DeepSearch 0.69, Grok3 DeeperSearch 0.80, Perplexity Deep Research 0.85; groundedness (cited claims / all factual claims): 0.34, 0.59, 0.32, 0.31, and 0.56 respectively","comparison":"high conditional citation support coexisted with lower citation coverage, including OpenAI Deep Research 0.84 faithfulness and 0.34 groundedness"},"scope":{"models_or_agents":["OpenAI Deep Research","Gemini Deep Research powered by Gemini-2.5-Pro","Grok3 DeepSearch","Grok3 DeeperSearch","Perplexity Deep Research"],"dataset_or_population":"65 expert-selected frontier-AI research questions across 35 AI subjects and technical-detail, literature-review, and open-consulting types","tool_and_retrieval_path":"section-level claim extraction and source-text extraction; GPT-4.1 was selected as extraction and factual-assessment judge","time":"all system evaluations conducted March-April 2025","metric_scope":"faithfulness conditions on cited claims and asks claim support; groundedness is citation coverage, not a claim-truth rate and not a URL-liveness measure"},"source_ids":["s3_researcherbench"],"exact_span_ids":["s3_definition","s3_models_time","s3_table"],"interpretation":"Verified study result: its conditional support metric cannot be read as the share of all report claims that are correct or supported."},{"result_id":"r4_cited_not_verified_source_attribution","proposition":"Cited but Not Verified jointly evaluated link accessibility, topical relevance, and factual claim support against retrieved cited content.","reported_value":{"numerator":"unknown factual-support counts; scores are proportions over evaluated attributions","denominator":"130 research queries; per-model attribution-pair denominators unknown in the reported table","rate":"14-model table Fact Check ranged from 24.4% to 76.8%; named frontier examples: Claude Opus 4.5 76.8%, GPT-5.4 47.7%, GPT-5.2 58.8%, Codex 54.1%, Claude Haiku 4.5 68.9%, GPT-5 Mini 38.9%, Gemini 3 Flash 45.2%, Gemini 3.1 Pro 48.5%; Link Works was 94.1%-100.0% for listed frontier models and Relevant Content 80.7%-95.7%","comparison":"at 2 versus 150 tool calls, GPT-5.4 Fact Check fell 78.6% to 16.7%; Claude Opus 4.6 fell 80.0% to 57.9%; the paper reports an approximately 42% average decline while Link Works and Relevant Content stayed above 92% at all tested depths"},"scope":{"models_or_agents":["14 closed-source and open-source LLMs, including Claude, GPT, Gemini, Codex, OSS-120B, Pixtral Large, and Llama 4 Maverick"],"dataset_or_population":"130 research queries drawn from DeepResearch Bench; cited Markdown reports that could be parsed into citation-claim pairs","tool_and_retrieval_path":"reproducible Markdown AST parser; retrieval of cited content; rubric-based LLM-as-a-judge evaluators calibrated through manual review of 50-100 judgments","time":"preprint posted 2026-05-07; output-generation dates are not fully specified in the verified spans","metric_scope":"Link Works is accessibility, Relevant Content is topical alignment, and Fact Check is binary support/consistency for facts, numbers, dates, and assertions; the agent/model list is broader than products marketed as dedicated deep-research agents"},"source_ids":["s4_cited_not_verified"],"exact_span_ids":["s4_definition_table","s4_depth"],"interpretation":"Verified study result: high link and relevance scores did not imply high claim-support scores in this study. It is not independent evidence about the task set from DeepResearch Bench, because its 130 queries were drawn from that benchmark."},{"result_id":"r5_urlhealth_resolution","proposition":"Detecting and Correcting Reference Hallucinations measured whether citation URLs resolved and operationally classified non-resolving URLs as archived/stale or absent from the Wayback Machine/hallucinated.","reported_value":{"numerator":"OpenAI Deep Research: 416 hallucinated-equivalent URLs and approximately 416 stale URLs implied by 3.5% and 6.6% of 4,121; Gemini-2.5-Pro Deep Research: approximately 1,504 hallucinated and approximately 588 stale URLs implied by 13.3% and 5.2% of 11,309; exact raw counts are not printed in the verified table","denominator":"DRBench: 4,121 URLs from OpenAI Deep Research and 11,309 from Gemini-2.5-Pro Deep Research; across both study datasets: 53,090 DRBench URLs and 168,021 ExpertQA URLs","rate":"OpenAI Deep Research 10.1% non-resolving and 3.5% no-archive/hallucinated; Gemini-2.5-Pro Deep Research 18.5% non-resolving and 13.3% no-archive/hallucinated","comparison":"pooled two deep-research agents: 10.7% hallucinated versus 4.8% for eight search-augmented models, z=15.15, p<10^-51; non-resolving 16.2% versus 6.8%, z=20.20, p<10^-89"},"scope":{"models_or_agents":["OpenAI Deep Research","Gemini-2.5-Pro Deep Research","eight search-augmented models on DRBench; three search-augmented models on ExpertQA"],"dataset_or_population":"DRBench 100 multilingual research queries with pre-collected model outputs, plus ExpertQA 2,177 expert-curated questions across 32 fields","tool_and_retrieval_path":"regex/API URL extraction; HTTP HEAD with GET fallback; 4xx/5xx, connection failure, or timeout classified non-resolving except 403; Wayback Machine lookup classified non-resolving URLs with no snapshot as hallucinated and those with a snapshot as stale","time":"preprint v1 2026-04-03; the DRBench outputs are reused benchmark outputs rather than a newly sampled population","metric_scope":"URL accessibility and archival presence only; the paper expressly does not systematically measure whether source content supports the associated claim"},"source_ids":["s5_urlhealth"],"exact_span_ids":["s5_scope","s5_table","s5_comparison"],"interpretation":"Verified study result: this is direct evidence about resolvability, not semantic entailment. It should therefore not be combined arithmetically with support rates from other studies."}],"sources":[{"source_id":"s1_deepresearchbench","url":"https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf","title":"DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents","authors_or_org":"Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao","date":"2025-06-13","identifier":"arXiv:2506.11763; later published at ICLR 2026","edition":"31-page public PDF examined; title page labels it \"Preprint. Work in progress.\"","retrieval_status":"verified_public_pdf","media_type":"PDF research preprint/conference version","license":"unknown","exact_spans":[{"span_id":"s1_method","locator":"p. 4, lines 150-159","quote":"\"Each unique Statement-URL pair undergoes a support evaluation... This results in a binary judgment ('support' or 'not support') for each pair, determining whether the citation accurately grounds the claim.\"","supports":"FACT directly evaluates claim-to-citation support."},{"span_id":"s1_table","locator":"p. 5, Table 1, lines 179-198","quote":"\"Perplexity Deep Research ... 90.24 31.26\"; \"Gemini-2.5-Pro Deep Research ... 81.44 111.21\"; \"OpenAI Deep Research ... 77.96 40.79\".","supports":"Reported FACT Citation Accuracy and Effective Citations values."},{"span_id":"s1_dates","locator":"p. 17, Table 5, lines 715-729","quote":"\"OpenAI Deep Research April 1 – May 8\"; \"Gemini 2.5 Pro Deep Research April 27 – April 29\"; \"Perplexity Deep Research April 1 – April 29\".","supports":"Output-vintage scope."}]},{"source_id":"s2_reportbench","url":"https://arxiv.org/pdf/2508.15804","title":"ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks","authors_or_org":"Minghao Li, Ying Zeng, Zhihao Cheng, Cong Ma, Kai Jia; ByteDance BandAI","date":"2025-08-14","identifier":"arXiv:2508.15804v1","edition":"18-page arXiv v1 PDF examined","retrieval_status":"verified_public_pdf","media_type":"PDF research preprint","license":"unknown","exact_spans":[{"span_id":"s2_method","locator":"p. 5, lines 289-300","quote":"\"we retrieve the full content of each cited webpage via web scraping... [and] perform consistency verification by comparing the statement with the retrieved content\".","supports":"Citation-claim consistency method."},{"span_id":"s2_metric_table","locator":"p. 6, Table 1, lines 333-355","quote":"\"For cited statements, we compute the match rate, i.e., the proportion of statements that are semantically consistent with their cited sources.\"; \"OpenAI Deep Research ... 78.87% 88.2\"; \"Gemini Deep Research ... 72.94% 96.2\".","supports":"Metric definition and values."},{"span_id":"s2_collection","locator":"p. 5, lines 315-329","quote":"\"we manually collected responses ... during the period from July 14 to July 25... OpenAI was using the standard version of Deep Research, powered by the o3 model.\"","supports":"Agent configuration and collection period."}]},{"source_id":"s3_researcherbench","url":"https://arxiv.org/pdf/2507.16280","title":"ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry","authors_or_org":"ResearcherBench authors; SII-GAIR","date":"2025-07-22","identifier":"arXiv:2507.16280v1","edition":"22-page arXiv v1 PDF examined","retrieval_status":"verified_public_pdf","media_type":"PDF research preprint","license":"unknown","exact_spans":[{"span_id":"s3_definition","locator":"p. 6, lines 361-387","quote":"\"Faithfulness score... evaluates the proportion of cited claims that are actually supported by their referenced sources\"; \"Groundedness score... measures the proportion of all factual claims that have explicit citation support\".","supports":"Support and coverage definitions."},{"span_id":"s3_models_time","locator":"p. 6-7, lines 395-409","quote":"\"We evaluated several leading commercial deep research systems\" including OpenAI, Gemini, Grok, and Perplexity; \"All evaluations were conducted between March and April in 2025\".","supports":"System and time scope."},{"span_id":"s3_table","locator":"p. 7, Table 2, lines 415-427","quote":"\"OpenAI Deep Research 0.7032 0.84 0.34\"; \"Gemini Deep Research 0.6929 0.86 0.59\"; \"Perplexity Deep Research 0.4800 0.85 0.56\".","supports":"Faithfulness and groundedness results."}]},{"source_id":"s4_cited_not_verified","url":"https://arxiv.org/pdf/2605.06635","title":"Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents","authors_or_org":"Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah, Corey Feld","date":"2026-05-07","identifier":"arXiv:2605.06635","edition":"12-page arXiv PDF examined","retrieval_status":"verified_public_pdf","media_type":"PDF research preprint","license":"unknown","exact_spans":[{"span_id":"s4_definition_table","locator":"pp. 5-6, lines 209-217 and Table 1, lines 254-271","quote":"\"Fact Check verifies whether specific factual claims are accurately supported by the source content\"; \"Claude Opus 4.5 ... 98.7% 95.7% 76.8%\"; \"GPT-5 Mini ... 99.3% 87.4% 38.9%\".","supports":"Support metric, evaluation protocol, and model results."},{"span_id":"s4_depth","locator":"pp. 6-7, lines 286-318","quote":"\"Fact Check accuracy drops approximately 42% on average\"; \"GPT-5.4 ... 78.6%\" at 2 calls and \"16.7%\" at 150; Claude Opus 4.6 \"80.0%\" at 2 and \"57.9%\" at 150.","supports":"Search-depth ablation."}]},{"source_id":"s5_urlhealth","url":"https://arxiv.org/pdf/2604.03173","title":"Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents","authors_or_org":"Delip Rao, Eric Wong, Chris Callison-Burch","date":"2026-04-03","identifier":"arXiv:2604.03173v1","edition":"25-page arXiv v1 PDF examined","retrieval_status":"verified_public_pdf","media_type":"PDF research preprint","license":"unknown","exact_spans":[{"span_id":"s5_scope","locator":"pp. 1-2, lines 64-95","quote":"\"This study focuses on URL-based citation hallucinations\"; \"fabricated snippets ... and invented bibliographic entries ... require separate systematic study.\"","supports":"The study does not measure semantic support."},{"span_id":"s5_table","locator":"p. 3, Table 2, lines 173-200","quote":"\"openai-deepresearch OpenAI 4,121 10.1 ... 3.5 ... 6.6\"; \"gemini-2.5-pro-deepres. Google 11,309 18.5 ... 13.3 ... 5.2\".","supports":"Per-model resolution and Wayback-based classification rates."},{"span_id":"s5_comparison","locator":"p. 4, lines 210-226","quote":"\"Pooling across the two deep research agents, the hallucination rate is 10.7% ... versus 4.8% ... for the eight search-augmented models\"; \"Non-resolving rates ... 16.2% ... versus 6.8%\".","supports":"Pooled comparison and statistical basis."}]}],"counterevidence":[{"claim":"All evidence points to citation failure.","evidence":"No. Conditional support metrics are often high in particular settings: ResearcherBench reported 0.84-0.86 faithfulness for OpenAI and Gemini Deep Research, and DeepResearch Bench reported 77.96%-90.24% task-averaged citation accuracy for four deep-research agents.","source_ids":["s1_deepresearchbench","s3_researcherbench"],"qualification":"These are not conflicting measurements of the same runs or the same denominator, so they should not be averaged with ReportBench or Cited but Not Verified."},{"claim":"A resolving or relevant URL establishes that the claim is supported.","evidence":"No. Cited but Not Verified reported frontier Link Works above 94% and relevance above 80%, while Fact Check ranged from 39% to 77% in its headline result and 24.4%-76.8% in its complete table.","source_ids":["s4_cited_not_verified"],"qualification":"Its support decisions are LLM-judge results calibrated by limited manual review, not universal ground truth."},{"claim":"URL-hallucination rates measure semantic claim support.","evidence":"No. urlhealth explicitly limits scope to URL-based citation hallucinations and says fabricated snippets and invented bibliographic entries require a separate systematic study.","source_ids":["s5_urlhealth"],"qualification":"Its URL rates are resolution/archival evidence only."}],"limitations":["All reported support measures use automated LLM judgment or semantic matching; DeepResearch Bench reports a 100-pair human comparison for its selected judge, and Cited but Not Verified reports calibration through manual review of 50-100 judgments, but neither is a full human audit of every reported pair.","Different studies use different units: unique statement-URL pairs, cited statements, cited claims, all factual claims, URLs, task means, and report means. Their rates are not a common numerator/denominator.","DeepResearch Bench commercial-output data are reused by urlhealth; urlhealth therefore adds a different failure measurement on overlapping outputs, not an independent claim-support replication. Cited but Not Verified draws 130 queries from DeepResearch Bench, so its task population is also not wholly independent.","URL liveness is time-sensitive. urlhealth's no-Wayback-snapshot criterion is operational, and the paper notes incomplete and nonuniform archive coverage; it also excludes HTTP 403 responses, which can be live but bot-blocked.","Web retrieval can fail on paywalls, bot restrictions, changing pages, malformed citation formatting, and parser limitations. A retrieved page can be real and topically relevant yet not contain the particular asserted fact.","The evaluated commercial systems are dated snapshots or interface configurations, not stable product identities. Results must not be generalized to later versions, different runtime settings, private-source modes, other prompts, or all agents.","ReportBench's non-cited-statement factual-accuracy result is outside the citation-support question and is not used here as evidence that citations supported claims."],"unresolved":["A public primary manual dermatology evaluation exists: Keplinger, Frashure, Duran, and Hu, Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype, DOI 10.1111/jdv.70035, first published 2025-09-04. Its public Mendeley dataset (DOI 10.17632/3s73z9zf3c.2, published 2025-07-21) lists prompts, evaluation results, and unedited outputs under CC BY 4.0. The accessible public record did not expose the result files or article full text in this trace, so no reported quantitative claim from it is included as verified evidence.","Exact aggregate supported-pair numerators and denominators are unavailable in the cited table-level spans for DeepResearch Bench, ReportBench, ResearcherBench, and Cited but Not Verified. Rates are reported, but reconstructing raw counts would require their released artifacts and may not be exact because several metrics average per task or report.","No verified source in this packet validates whether the claims themselves are true independent of their cited sources, beyond the studies' cited-source support/consistency definitions."],"search_notes":["Used public, credential-free primary papers and a public official data artifact only; excluded journalistic summaries, vendor marketing claims, search-result snippets, generated summaries, and unverified secondary quotations.","Included separate URL-resolution evidence because the question asks whether citations resolve; labeled it distinctly from claim-support evidence.","Search terms targeted deep-research-agent citation accuracy, citation faithfulness, source attribution, URL validity, ReportBench, DeepResearch Bench, ResearcherBench, and the dermatology primary evaluation. The findings are a bounded high-relevance packet, not a proof that no other public study existed by the cutoff."]}

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0