Repository object · research-note
V2 Terra 04
Accepted research note in the public catalog.
- Media type
application/json- Object ID
em:research-note:sha256:521cd682380fe4f2dcc48d4f53f6aff6a55f106e85c48a75cab594bdd63a7aa4- Content digest
8deffa91b6bb52d1de4ef4965d0087eca3d3a8a96d971a5e89753c1928cc4ef4
Also filed under
Source content
{"question":"What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?","cutoff":"2026-08-22","answer":"Public primary research by the cutoff contains direct measurements of both dimensions, but they measure different agent versions, task populations, and units. URL resolution was directly measured at scale in Rao, Wong, and Callison-Burch (2026): for the two deep-research agents in its DRBench subset, OpenAI Deep Research had 10.1% non-resolving URLs (3.5% classified hallucinated; 4,121 URLs) and Gemini 2.5 Pro Deep Research had 18.5% non-resolving (13.3% hallucinated; 11,309 URLs). Claim-source support was directly measured by four public benchmarks: DeepResearch Bench reported 77.96%/81.44%/90.24%/83.59% citation accuracy for OpenAI/Gemini/Perplexity/Grok respectively; ResearcherBench reported faithfulness 0.84/0.86/0.85/0.69-0.80 for those systems; ReportBench reported 78.87% and 72.94% cited-statement match rates for OpenAI and Gemini; and Onweller et al. reported Fact Check rates from 24.4% to 76.8% across 14 later web-search research-agent configurations despite generally high URL accessibility and topical relevance. These are not commensurate estimates and must not be pooled. The strongest direct evidence therefore supports neither universal success nor universal failure: cited URLs often resolved in the studied settings, but support/faithfulness varied materially by system, study, metric, task population, and time.","results":[{"result_id":"R1_deepresearchbench_support","proposition":"In DeepResearch Bench's FACT evaluation, a cited statement-URL pair was counted as accurate only when retrieved page text supported the claim; citation accuracy varied across four commercial deep-research agents.","reported_value":{"OpenAI_Deep_Research":{"numerator":"unknown","denominator":"100 task-level accuracies; raw supported-pair counts unknown","rate":"77.96% citation accuracy","comparison":"40.79 average effective citations per task"},"Gemini_2_5_Pro_Deep_Research":{"numerator":"unknown","denominator":"100 task-level accuracies; raw supported-pair counts unknown","rate":"81.44% citation accuracy","comparison":"111.21 average effective citations per task"},"Perplexity_Deep_Research":{"numerator":"unknown","denominator":"100 task-level accuracies; raw supported-pair counts unknown","rate":"90.24% citation accuracy","comparison":"31.26 average effective citations per task"},"Grok_Deeper_Search":{"numerator":"unknown","denominator":"100 task-level accuracies; raw supported-pair counts unknown","rate":"83.59% citation accuracy","comparison":"8.15 average effective citations per task"},"quantitative_basis":"The study extracts and deduplicates statement-URL pairs, retrieves webpage text through Jina Reader API, uses Gemini-2.5-Flash for binary support judgments, calculates each task's supported-pair proportion, assigns zero for no citable statement, then averages over tasks."},"scope":{"models_or_agents":["OpenAI Deep Research","Gemini-2.5-Pro Deep Research","Perplexity Deep Research","Grok Deeper Search"],"dataset":"DeepResearch Bench, 100 PhD-level tasks across 22 fields; 50 Chinese and 50 English","tool_or_evaluator":"Jina Reader API for webpage text; Gemini-2.5-Flash for statement-URL extraction and support judgment","population":"Deduplicated unique statement-URL pairs in generated reports; raw pair total unknown","time":"Commercial-agent outputs: OpenAI 2025-04-01 to 2025-05-08; Gemini 2025-04-27 to 2025-04-29; Perplexity 2025-04-01 to 2025-04-29; Grok 2025-04-27 to 2025-04-29","metric_scope":"Citation support, not a separately reported HTTP URL-resolution rate; macro-average of per-task support ratios, not necessarily a pooled pair-level percentage."},"source_ids":["S1_deepresearchbench"],"exact_span_ids":["S1_method","S1_table1","S1_metric"],"failure_classes":["irrelevant source","real source supporting a weaker claim","inaccessible source"],"interpretation":"This is direct support evidence, not proof that all URLs resolve. Its percentages are evaluation-model judgments over the benchmark and should not be treated as universal rates or raw supported-citation fractions."},{"result_id":"R2_researcherbench_support","proposition":"ResearcherBench measured whether cited claims were supported by text extracted from their cited URLs, reporting high but non-perfect faithfulness alongside much lower citation coverage.","reported_value":{"OpenAI_Deep_Research":{"numerator":"unknown","denominator":"cited claims; raw claim counts unknown","rate":"0.84 faithfulness","comparison":"0.34 groundedness"},"Gemini_Deep_Research":{"numerator":"unknown","denominator":"cited claims; raw claim counts unknown","rate":"0.86 faithfulness","comparison":"0.59 groundedness"},"Grok3_DeepSearch":{"numerator":"unknown","denominator":"cited claims; raw claim counts unknown","rate":"0.69 faithfulness","comparison":"0.32 groundedness"},"Grok3_DeeperSearch":{"numerator":"unknown","denominator":"cited claims; raw claim counts unknown","rate":"0.80 faithfulness","comparison":"0.31 groundedness"},"Perplexity_Deep_Research":{"numerator":"unknown","denominator":"cited claims; raw claim counts unknown","rate":"0.85 faithfulness","comparison":"0.56 groundedness"},"quantitative_basis":"Faithfulness = supported cited claims divided by cited claims; groundedness = cited claims divided by all extracted factual claims."},"scope":{"models_or_agents":["OpenAI Deep Research","Gemini Deep Research powered by Gemini-2.5-Pro","Grok3 DeepSearch","Grok3 DeeperSearch","Perplexity Deep Research"],"dataset":"ResearcherBench: 65 frontier-AI research questions across 35 AI subjects; 12 technical-detail, 20 literature-review, and 33 open-consulting questions","tool_or_evaluator":"Jina Reader API to extract cited URL text; GPT-4.1 as claim extractor and binary factual-support judge","population":"Extracted URL-claim-context triplets from reports; raw cited-claim totals unknown","time":"All evaluations March-April 2025; OpenAI collection 2025-03-24 to 2025-04-29, Perplexity 2025-03-24 to 2025-04-15, Gemini 2025-04-15 to 2025-04-21, and Grok variants late March to April 2025","metric_scope":"Claim-source support for claims bearing an extracted citation URL; it does not separately report URL HTTP resolution or test non-cited claims for source support."},"source_ids":["S2_researcherbench"],"exact_span_ids":["S2_method","S2_table2","S2_scope"],"failure_classes":["inaccessible source","irrelevant source","real source supporting a weaker claim"],"interpretation":"This independently reports claim-support ratios, but only in frontier-AI questions and with a GPT-4.1 judge. Its results cannot be pooled with DeepResearch Bench even where agent labels overlap."},{"result_id":"R3_reportbench_support","proposition":"ReportBench's cited-statement pipeline retrieved cited webpages, located a semantically relevant passage, and checked consistency; OpenAI Deep Research and Gemini Deep Research had reported cited-statement match rates below 100%.","reported_value":{"OpenAI_Deep_Research":{"numerator":"unknown","denominator":"cited statements; raw count unavailable because 88.2 is the average cited-statement count per report","rate":"78.87% match rate","comparison":"9.89 references and 88.2 cited statements per report"},"Gemini_Deep_Research":{"numerator":"unknown","denominator":"cited statements; raw count unavailable because 96.2 is the average cited-statement count per report","rate":"72.94% match rate","comparison":"32.42 references and 96.2 cited statements per report"},"quantitative_basis":"Match rate is the proportion of cited statements semantically consistent with cited sources after webpage retrieval and an LLM-based supporting-passage/consistency pipeline."},"scope":{"models_or_agents":["OpenAI Deep Research standard version powered by o3","Gemini Deep Research with Gemini 2.5 Pro and Deep Research toggles enabled"],"dataset":"ReportBench, 100 reverse-engineered research tasks based on expert-written arXiv survey papers; predominantly STEM","tool_or_evaluator":"Web scraping; GPT-4o for statement extraction, supporting-source extraction, and semantic-consistency verification","population":"Explicitly linked cited statements in generated academic-survey reports; raw statement-pair total unknown","time":"Responses manually collected from product web interfaces 2025-07-14 to 2025-07-25","metric_scope":"Cited-statement semantic consistency, not direct HTTP liveness; reference precision/recall is a separate metric and does not establish claim support."},"source_ids":["S3_reportbench"],"exact_span_ids":["S3_method","S3_table1","S3_time"],"failure_classes":["inaccessible source","irrelevant source","real source supporting a weaker claim"],"interpretation":"This is direct support evidence for an academic-survey task population. It is methodologically distinct from S1 and S2, and its score should not be described as an all-web or all-agent citation-accuracy rate."},{"result_id":"R4_urlhealth_resolution","proposition":"A large-scale URL-liveness study directly measured non-resolving and likely-nonexistent citation URLs from two deep-research agents in DRBench.","reported_value":{"OpenAI_Deep_Research":{"numerator":"unknown","denominator":"4,121 extracted URLs","rate":"10.1% non-resolving [9.1,11.0]; 3.5% hallucinated [3.0,4.1]; 6.6% stale","comparison":"41.2 URLs per query"},"Gemini_2_5_Pro_Deep_Research":{"numerator":"unknown","denominator":"11,309 extracted URLs","rate":"18.5% non-resolving [17.8,19.2]; 13.3% hallucinated [12.7,13.9]; 5.2% stale","comparison":"113.1 URLs per query"},"pooled_deep_research_agents":{"numerator":"unknown","denominator":"15,430 URLs if the two listed denominators are summed; the source's pooled numerator is not reported","rate":"10.7% hallucinated [10.2,11.2] and 16.2% non-resolving [15.7,16.8]","comparison":"search-augmented-model comparison: 4.8% hallucinated and 6.8% non-resolving"},"quantitative_basis":"URL HEAD request with GET fallback; 4xx/5xx, connection errors, and timeouts classified non-resolving except 403; non-resolving URLs with no Wayback snapshot classified hallucinated, otherwise stale."},"scope":{"models_or_agents":["openai-deepresearch","gemini-2.5-pro-deepresearch"],"dataset":"DRBench/DeepResearch Bench outputs: 100 multilingual research queries in finance, science, and technology; this study analyzed 10 of 23 available models where liveness data was reliable","tool_or_evaluator":"HTTP HEAD/GET liveness checks, browser-like User-Agent, Wayback Machine API","population":"Extracted URL citations, not claim-URL pairs","time":"Study publicly posted 2026-04-03; original DRBench output generation dates are not established by this source's result table","metric_scope":"URL resolution/existence only. It does not measure whether a resolving URL supports the associated claim."},"source_ids":["S4_urlhealth"],"exact_span_ids":["S4_table2","S4_comparison","S4_method"],"failure_classes":["nonexistent URL","non-resolving URL"],"interpretation":"This is the strongest direct URL-resolution evidence located. It is not support/entailment evidence. Its DRBench source outputs overlap the DeepResearch Bench artifact family, so S1 and S4 are complementary analyses of a related underlying output corpus, not independent replications."},{"result_id":"R5_cited_not_verified_support_and_resolution","proposition":"A later end-to-end study of citation-claim pairs found high working-link and topical-relevance rates but materially lower Fact Check support rates across 14 web-search research-agent configurations.","reported_value":{"range_across_14_models":{"numerator":"unknown","denominator":"Citation-claim evaluations; raw totals not reported in Table 1","rate":"Link Works 80.8%-100.0%; Relevant Content 60.6%-95.7%; Fact Check 24.4%-76.8%","comparison":"12 of 14 models exceeded 94% Link Works; frontier models exceeded 80% Relevant Content"},"selected_models":{"Claude_Opus_4_5":{"Link_Works":"98.7%","Relevant_Content":"95.7%","Fact_Check":"76.8%","query_success":"90.0%"},"GPT_5_4":{"Link_Works":"100.0%","Relevant_Content":"93.7%","Fact_Check":"47.7%","query_success":"100.0%"},"Gemini_3_1_Pro":{"Link_Works":"94.1%","Relevant_Content":"80.7%","Fact_Check":"48.5%","query_success":"90.0%"},"OSS_120B":{"Link_Works":"83.9%","Relevant_Content":"68.7%","Fact_Check":"24.4%","query_success":"40.0%"}},"quantitative_basis":"A deterministic Markdown AST parser extracts citation-claim pairs; each cited URL is fetched; Link Works is a binary accessibility check; Relevant Content and Fact Check are calibrated LLM-judge binary scores."},"scope":{"models_or_agents":["GPT-5.2","GPT-5.4","GPT-5 Mini","Codex","Claude Sonnet 4.5","Claude Sonnet 4.6","Claude Opus 4.5","Claude Opus 4.6","Claude Haiku 4.5","Gemini 3.1 Pro","Gemini 3 Flash","OSS-120B","Llama 4 Maverick","Pixtral Large"],"dataset":"130 research queries drawn from DeepResearch Bench and BrowseComp","tool_or_evaluator":"Markdown AST parser; source fetching; LLM-as-a-judge for relevance and factual support, calibrated through manual review of 50-100 judgments","population":"Citation-claim pairs in one-shot Markdown reports generated by web-search-enabled agents; the system prompt required citation format and minimum search depth","time":"Publicly posted 2026-05-07; model/runtime versions are those named in the paper and not equivalent to the 2025 product evaluations","metric_scope":"Directly measures both accessibility and claim-source support, but Fact Check is an LLM-judge result against retrieved content, not a human-adjudicated ground truth and not a commercial Deep Research product leaderboard."},"source_ids":["S5_cited_not_verified"],"exact_span_ids":["S5_method","S5_table1","S5_scope"],"failure_classes":["non-resolving URL","inaccessible source","irrelevant source","real source supporting a weaker claim"],"interpretation":"The result directly demonstrates that a resolving, topically related citation is not sufficient evidence of claim support in this experimental setting. It does not justify a claim about every deployed deep-research product."},{"result_id":"R6_search_depth_ablation","proposition":"In the same later study, increasing allowed search-tool calls lowered Fact Check support while Link Works and relevance stayed high for two frontier model-agent configurations.","reported_value":{"GPT_5_4":{"numerator":"unknown","denominator":"citation-claim evaluations at each tool-call cap; raw totals unknown","rate":"Fact Check 78.6% at 2 calls and 16.7% at 150 calls","comparison":"Link Works 100.0% and 99.2%; Relevant Content 100.0% and 99.2% at those endpoints"},"Claude_Opus_4_6":{"numerator":"unknown","denominator":"citation-claim evaluations at each tool-call cap; raw totals unknown","rate":"Fact Check 80.0% at 2 calls and 57.9% at 150 calls","comparison":"Link Works 100.0% at both endpoints; Relevant Content 100.0% at both endpoints"},"combined_statement":{"numerator":"unknown","denominator":"two model configurations","rate":"approximately 42% average Fact Check decrease from 2 to 150 calls","comparison":"surface metrics stayed above 92% at all tested depths"},"quantitative_basis":"Controlled ablation at 2, 10, 30, 50, 70, 100, and 150 maximum tool calls."},"scope":{"models_or_agents":["GPT-5.4","Claude Opus 4.6"],"dataset":"The S5 study's 130-query DeepResearch Bench/BrowseComp population; per-depth query/pair totals unknown","tool_or_evaluator":"Same AST extraction, source retrieval, Link Works, Relevant Content, and Fact Check pipeline as S5","population":"Citation-claim pairs generated under fixed maximum search-depth caps","time":"Publicly posted 2026-05-07","metric_scope":"A controlled tool-call-cap comparison within two configurations. It is not evidence that all deep-research agents or all retrieval strategies worsen with additional search."},"source_ids":["S5_cited_not_verified"],"exact_span_ids":["S5_depth"],"failure_classes":["real source supporting a weaker claim"],"interpretation":"This is a within-study counterpoint to the idea that more retrieval necessarily improves citation support, but it is not independent evidence from R5 and should not be counted as a separate replication."}],"sources":[{"source_id":"S1_deepresearchbench","url":"https://arxiv.org/html/2506.11763","title":"DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents","authors_or_org":"Mingxuan Du; Benfeng Xu; Chiwei Zhu; Xiaorui Wang; Zhendong Mao","date":"2025-06-13","identifier":"arXiv:2506.11763v1; DOI 10.48550/arXiv.2506.11763","edition":"arXiv v1, 31 pages","retrieval_status":"retrieved publicly without credentials on 2026-08-22","media_type":"primary research preprint and public benchmark artifact","license":"CC BY 4.0","exact_spans":[{"span_id":"S1_method","locator":"Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment'","quote":"\"Each unique Statement-URL pair undergoes a support evaluation.\"","supports":"The reported citation metric evaluates an extracted claim against retrieved cited-page content."},{"span_id":"S1_table1","locator":"Table 1, Deep Research Agent rows","quote":"\"Perplexity Deep Research ... 90.24 ... 31.26\"","supports":"Perplexity's FACT citation accuracy and effective-citation result; the adjacent rows provide Gemini, OpenAI, and Grok values."},{"span_id":"S1_metric","locator":"Appendix E.1, equations (4)-(5)","quote":"\"C. Acc. assesses the ... proportion of 'support' statement-URL pairs\"","supports":"The denominator and macro-averaging definition."},{"span_id":"S1_time","locator":"Appendix D, Table 5","quote":"\"OpenAI Deep Research | April 1 – May 8\"","supports":"Collection-time scope for the evaluated commercial outputs."}]},{"source_id":"S2_researcherbench","url":"https://arxiv.org/html/2507.16280","title":"ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry","authors_or_org":"Tianze Xu; Pengrui Lu; Lyumanshan Ye; Xiangkun Hu; Pengfei Liu","date":"2025-07-22","identifier":"arXiv:2507.16280v1; DOI 10.48550/arXiv.2507.16280","edition":"arXiv v1, 22 pages","retrieval_status":"retrieved publicly without credentials on 2026-08-22","media_type":"primary research preprint and public benchmark artifact","license":"arXiv.org perpetual non-exclusive license","exact_spans":[{"span_id":"S2_method","locator":"Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3)","quote":"\"whether the extracted content supports the corresponding claim\"","supports":"The study's faithfulness numerator is direct claim-source support."},{"span_id":"S2_table2","locator":"Table 2, Deep Research System rows","quote":"\"OpenAI Deep Research | 0.7032 | 0.84 | 0.34\"","supports":"OpenAI's coverage, faithfulness, and groundedness values; adjacent rows report Gemini, Grok, and Perplexity."},{"span_id":"S2_scope","locator":"Abstract and Section 5.1 'Evaluation Configuration'","quote":"\"65 research questions ... across 35 different AI subjects\"","supports":"Task-domain and population scope."},{"span_id":"S2_time","locator":"Section 5.1 and Appendix E, Table 4","quote":"\"All evaluations were conducted between March and April in 2025\"","supports":"Evaluation period and version-boundary caution."}]},{"source_id":"S3_reportbench","url":"https://arxiv.org/html/2508.15804","title":"ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks","authors_or_org":"Minghao Li; Ying Zeng; Zhihao Cheng; Cong Ma; Kai Jia; ByteDance BandAI","date":"2025-08-14","identifier":"arXiv:2508.15804v1; DOI 10.48550/arXiv.2508.15804","edition":"arXiv v1","retrieval_status":"retrieved publicly without credentials on 2026-08-22","media_type":"primary research preprint and public benchmark artifact","license":"CC BY 4.0","exact_spans":[{"span_id":"S3_method","locator":"Section 2.2 'Cited statements'","quote":"\"retrieve the full content of each cited webpage\"","supports":"The report measures cited-statement support against retrieved cited content."},{"span_id":"S3_table1","locator":"Section 3.2, Table 1","quote":"\"OpenAI Deep Research ... 78.87% ... 88.2\"","supports":"OpenAI match rate and average cited-statement count; Gemini's adjacent row gives 72.94% and 96.2."},{"span_id":"S3_time","locator":"Section 3.1 'Setttings'","quote":"\"during the period from July 14 to July 25\"","supports":"Data-collection dates and WebUI setup."},{"span_id":"S3_limitations","locator":"Appendix A.1 'Limitations'","quote":"\"most of which are concentrated in STEM fields\"","supports":"Dataset-domain limitation."}]},{"source_id":"S4_urlhealth","url":"https://arxiv.org/html/2604.03173","title":"Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents","authors_or_org":"Delip Rao; Eric Wong; Chris Callison-Burch; University of Pennsylvania","date":"2026-04-03","identifier":"arXiv:2604.03173v1; DOI 10.48550/arXiv.2604.03173","edition":"arXiv v1","retrieval_status":"retrieved publicly without credentials on 2026-08-22","media_type":"primary research preprint, public data, and official project artifact","license":"CC0","exact_spans":[{"span_id":"S4_table2","locator":"Section 4.1, Figure 1/Table 2","quote":"\"openai-deepresearch | OpenAI | 4,121 | 10.1 ... | 3.5 ... | 6.6\"","supports":"OpenAI Deep Research URL denominator and non-resolving/hallucinated/stale results; the Gemini row is adjacent."},{"span_id":"S4_comparison","locator":"Section 4.2, paragraph beginning 'Deep research agents cite far more URLs per query'","quote":"\"16.2% ... versus 6.8%\"","supports":"Pooled deep-research versus search-augmented non-resolution comparison."},{"span_id":"S4_method","locator":"Section 3.3 'URL extraction and classification'","quote":"\"URLs returning 4xx or 5xx status codes, connection errors, or timeouts are classified as non-resolving\"","supports":"Operational definition of URL resolution and Wayback-based stale/hallucinated split."},{"span_id":"S4_limitations","locator":"Section 2 definition discussion and Section 3.3","quote":"\"our non-resolving rates are lower bounds\"","supports":"403/bot-blocking and classification limitation."}]},{"source_id":"S5_cited_not_verified","url":"https://arxiv.org/html/2605.06635","title":"Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents","authors_or_org":"Hailey Onweller; Elias Lumer; Austin Huber; Pia Ramchandani; Vamse Kumar Subbiah; Corey Feld; PricewaterhouseCoopers U.S. Commercial Technology and Innovation Office","date":"2026-05-07","identifier":"arXiv:2605.06635v1; DOI 10.48550/arXiv.2605.06635","edition":"arXiv v1","retrieval_status":"retrieved publicly without credentials on 2026-08-22","media_type":"primary research preprint","license":"CC BY 4.0","exact_spans":[{"span_id":"S5_method","locator":"Section 3.3.1-3.3.3","quote":"\"Fact Check verifies whether specific factual claims are accurately supported by the source content.\"","supports":"Direct support criterion and distinction from URL accessibility/relevance."},{"span_id":"S5_scope","locator":"Section 3.4 'Experimental setup'","quote":"\"14 LLMs ... on 130 research queries drawn from DeepResearch Bench and BrowseComp\"","supports":"Model and dataset scope."},{"span_id":"S5_table1","locator":"Section 4.1, Table 1","quote":"\"Claude Opus 4.5 | 90.0% | 98.7% | 95.7% | 76.8%\"","supports":"Success, accessibility, relevance, and Fact Check rates; adjacent rows report the remaining models."},{"span_id":"S5_depth","locator":"Section 4.3, Tables 2-3 and following paragraph","quote":"\"Fact Check accuracy drops approximately 42% on average\"","supports":"Search-depth ablation summary."},{"span_id":"S5_limitations","locator":"Section 5 'Limitations'","quote":"\"LLM-as-a-judge ... may retain biases inherent to the judge model\"","supports":"Judge-based support-evaluation limitation."}]}],"counterevidence":[{"claim":"Citations from deep-research agents consistently support claims.","evidence_against":["S1 reports 77.96%-90.24% citation accuracy across four 2025 commercial agents, rather than 100%.","S3 reports 78.87% and 72.94% cited-statement match rates for its OpenAI and Gemini runs.","S5 reports Fact Check rates of 24.4%-76.8% despite often high link and relevance scores."]},{"claim":"Citations from deep-research agents generally fail or do not support claims.","evidence_against":["S1 reports citation-support accuracy above 77% for all four tested agents in its specific FACT metric.","S2 reports faithfulness 0.80-0.86 for four of five named deep-research systems, with Grok3 DeepSearch at 0.69.","S4 reports 81.5%-89.9% resolving URLs for its two deep-research agents under its specified liveness protocol."]},{"claim":"A working URL establishes evidential support.","evidence_against":["S5 reports 94.1%-100.0% Link Works for its frontier models but Fact Check values of 38.9%-76.8%; its study expressly measures these as distinct dimensions."]}],"limitations":[{"issue":"Metric non-equivalence","detail":"S1 uses a macro-average of task-level support ratios; S2 faithfulness is supported cited claims/cited claims; S3 match rate uses a GPT-4o semantic-consistency pipeline; S5 Fact Check is a calibrated LLM-judge score; S4 measures URL liveness/existence only. They must not be pooled, averaged, or ranked as one common rate."},{"issue":"Overlap and dependence","detail":"S4 analyzes URLs from DRBench/DeepResearch Bench outputs, so S1 and S4 are analyses of a related output artifact, not independent evidence. S5 draws queries from DeepResearch Bench but generates later outputs with different models/settings; task overlap does not make results an independent task-population replication."},{"issue":"Resolution is not support","detail":"S4 does not test whether a resolving source supports a claim. Conversely, S1-S3 primarily assess retrieved-source support and do not publish a separate ordinary HTTP-resolution success rate."},{"issue":"Automated adjudication","detail":"S1 uses Gemini-2.5-Flash, S2 GPT-4.1, S3 GPT-4o, and S5 an LLM judge for substantive support judgments. S5 reports manual calibration of only 50-100 judgments; no result is equivalent to exhaustive human adjudication."},{"issue":"Dynamic web and access artifacts","detail":"A URL can fail automated retrieval because of bot blocks, paywalls, JavaScript, timeout, or later link rot. S4 excludes 403s and calls its non-resolution estimates lower bounds; absence from Wayback is an imperfect proxy for nonexistence."},{"issue":"Version and time boundaries","detail":"Commercial agents are black boxes and studies collected outputs in distinct 2025 or 2026 windows. Labels such as 'OpenAI Deep Research' do not establish that the underlying runtime is identical across S1-S4, and S5 evaluates named later model configurations rather than the 2025 product runs."},{"issue":"Population bounds","detail":"S1 mixes Chinese/English PhD-level tasks across 22 domains; S2 is 65 frontier-AI questions; S3 is 100 mostly-STEM academic survey tasks; S4 uses multilingual DRBench plus a separate ExpertQA URL-only population; S5 uses 130 DeepResearch Bench/BrowseComp research queries. None supports generalization to all agents, languages, domains, or high-stakes uses."}],"unresolved":[{"item":"Raw numerators for claim-support results","status":"unknown","detail":"The public result tables report rates and, in some cases, average counts, but not the total number of adjudicated statement-URL or claim-URL pairs needed to reconstruct exact supported/unsupported numerators."},{"item":"Human-ground-truth error rates","status":"unknown","detail":"The public papers do not provide exhaustive independent human adjudication of every support or non-support decision for their headline commercial-agent results."},{"item":"Same-run joint probability of resolution and support","status":"unknown","detail":"No located source reports, for the same commercial product runs, a joint breakdown such as resolving-and-supporting, resolving-but-unsupported, and non-resolving citation-claim pairs at product level."},{"item":"Causal faithfulness","status":"unknown","detail":"The reported metrics test whether a cited source supports a generated claim, not whether that source causally influenced generation rather than being added post hoc."},{"item":"Current product behavior at cutoff","status":"unknown","detail":"The studies establish behavior for dated collections, not a verified 2026-08-22 live performance level for any commercial agent."}],"search_notes":{"method":"Searched public web and arXiv for primary studies and official public benchmark/project artifacts specifically combining deep-research-agent outputs with citation URL accessibility, source retrieval, and claim-support verification. Excluded commentary, search snippets, user anecdotes, vendor marketing claims without empirical citation verification, papers posted after the cutoff, and citation-generation work that did not empirically test cited outputs.","included_sources_reason":"S1-S3 directly retrieve cited source content and test claim-source support for named deep-research systems. S4 is the strongest direct large-scale URL-resolution/existence study covering two named deep-research agents. S5 is the most direct publicly posted end-to-end study measuring accessibility, topical relevance, and factual support together.","official_artifacts_checked":["https://github.com/Ayanami0730/deep_research_bench","https://github.com/GAIR-NLP/ResearcherBench","https://github.com/ByteDance-BandAI/ReportBench"],"exclusions":["Benchmark scores for report quality or answer correctness alone, because they do not establish citation resolution or claim-source support.","Studies of generic LLM academic-reference fabrication unless they specifically evaluated web/deep-research agent citations.","Secondary explainers, leaderboard mirrors, and generated summaries."],"freshness":"All cited source editions were publicly available by 2026-08-22; each result remains bounded to its documented collection period and edition."}}
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0