How We Know · Case 002 · skeptical
When eight research agents agree, how many evidence roots are there?
What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?
Eight context-isolated public-by-design reports captured under one frozen prompt, plus the public source editions and exact spans they cited. This historical pilot does not estimate current or universal agent behavior.
skeptical finding
This pilot cannot estimate current agent citation reliability: thirty-four citation occurrences remain unresolved, twenty claims required correction or no credit, four warrant groups remain pending, and zero warrant roots are independently confirmed by the packet. Inspect the exact source span before relying on a polished cited answer.
What to do with that: Do not infer independent corroboration from eight agreeing reports: 34 citation occurrences remain unresolved, 20 claims lost credit, and no warrant root was independently confirmed.
The packet is a lineage audit, not a vendor ranking or a representative product test.
Lineage accounting
Agreement is not a vote count
Shared capture: Unknown provider and retrieval dependencies remain; all reports share the exact prompt and one bounded capture program, so run multiplicity gets zero automatic independence credit.
Warrant boundary: Unknown residual independence remains across task data, judge methods, retrieval, source, edition, span, derivation, and upstream-citation lineages.
Sentence x-ray
What the selected record actually supports
Open a sentence to inspect its exact work, edition, span, retrieval, digest, and license chain.
01 This report is a captured observation, not an independent evidence root.
Typed relation: dependence
EM-0026 deterministic agent-citation evidence ledger
Captured report V2-TERRA-04
{ "answer_bytes": 30189, "answer_path": "research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-04.json", "answer_sha256": "8deffa91b6bb52d1de4ef4965d0087eca3d3a8a96d971a5e89753c1928cc4ef4", "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_model_profile": "gpt-5.6-terra", "retrieval_infrastructure": "unknown", "run_id": "V2-TERRA-04", "status": "completed", "trace_path": "research/how-we-know/agent-citation-lineage/traces-v2/V2-TERRA-04.json" }
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:58db8115527aba46c1706f0fc3999e811f0352f5e855ad09e56c8688118d94b8
- Span digest
- sha256:c9198194e5a51ebafd593b2d7e2386afb6ee76460f01bf2958f54df04981c528
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
EM-0026 deterministic agent-citation evidence ledger
Shared capture lineage
{ "automatic_independence_credit": 0, "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_profiles": [ "gpt-5.6-sol", "gpt-5.6-terra" ], "retrieval_infrastructure": "unknown" }
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:e120064fee0bdabe15299cb865ff2f7df66473291a13b7a765a4a09d16223f5e
- Span digest
- sha256:92b558b23559a96a551de8acb2627b148688d21ede2ca3fb9966196c4675dab6
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
02 The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
Typed relation: qualification
EM-0026 deterministic agent-citation evidence ledger
Pending warrant warrant:citation-verifier-calibration
warrant:citation-verifier-calibration
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:755999bc2f4f9858e910f7436ae3e011343fb8b295c7ac25ebe37124b9fa9a79
- Span digest
- sha256:da47a19a65e8c1a6816fe474c99558e32442ba728c9fe0046c1ee24d79ec33d4
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
03 The complete unresolved set remains visible and receives no warrant credit.
Typed relation: undercutting
EM-0026 deterministic agent-citation evidence ledger
All unresolved citation occurrences
[ { "citation_occurrence_id": "V2-SOL-01:s1_keplinger", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1_keplinger", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 111993, "captured_sha256": "abbaa9f22035230d13152a68aaeecd8e83e40867c9dae4acd1d547e36c866793", "edition_id": "edition:keplinger-vor-2025", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolved_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "retrieval_status": "retrieved" }, "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-01:s1_keplinger:sp1a", "V2-SOL-01:s1_keplinger:sp1b", "V2-SOL-01:s1_keplinger:sp1c" ] }, { "citation_occurrence_id": "V2-SOL-01:s1a_keplinger_data", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v2", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans; file API required authentication and was not used", "raw_source_id": "s1a_keplinger_data", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 116951, "captured_sha256": "0a2c4c6ed54dd7e2e0ca6cc07aa19b18fa1439e506c2310db778094282d2cebc", "edition_id": "edition:keplinger-supplement-v2", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-01:s1a_keplinger_data:sp1d" ] }, { "citation_occurrence_id": "V2-SOL-01:s2_drbench", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s2_drbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 2665857, "captured_sha256": "8f80ce247f7cc355bb6773f36037e01cc3ba2e2082c77381eaadf0d2b92b021c", "edition_id": "edition:drbench-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-01:s2_drbench:sp2a", "V2-SOL-01:s2_drbench:sp2b", "V2-SOL-01:s2_drbench:sp2c", "V2-SOL-01:s2_drbench:sp2d" ] }, { "citation_occurrence_id": "V2-SOL-01:s3_deeptrace", "correction_ids": [], "edition_id": "edition:deeptrace-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_deeptrace", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 2982839, "captured_sha256": "dea4981c1066d0240a005f603b3b14419c7e32beb5191b3df847ceb26af3d6b6", "edition_id": "edition:deeptrace-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-SOL-01:s3_deeptrace:sp3a", "V2-SOL-01:s3_deeptrace:sp3b", "V2-SOL-01:s3_deeptrace:sp3c", "V2-SOL-01:s3_deeptrace:sp3d" ] }, { "citation_occurrence_id": "V2-SOL-01:s4_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 43848, "captured_sha256": "d1b6d476b3e81460dbc8821757775a8fa589e0579b77167ef0f33d35cb819cb5", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/abs/2605.06635", "resolved_url": "https://arxiv.org/abs/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/abs/2605.06635", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-SOL-01:s4_cited_not_verified:sp4a", "V2-SOL-01:s4_cited_not_verified:sp4b", "V2-SOL-01:s4_cited_not_verified:sp4c" ] }, { "citation_occurrence_id": "V2-SOL-01:s5_url_health", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s5_url_health", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 42408, "captured_sha256": "3976893e82f91cc9e7826c0d2d08c6caab40778134dbab09ebbfe1f6f63f5397", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/abs/2604.03173", "resolved_url": "https://arxiv.org/abs/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/abs/2604.03173", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-01:s5_url_health:sp5a", "V2-SOL-01:s5_url_health:sp5b", "V2-SOL-01:s5_url_health:sp5c" ] }, { "citation_occurrence_id": "V2-SOL-02:S1", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173v1", "resolved_url": "https://arxiv.org/html/2604.03173v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-02:S1:S1_SPAN_1", "V2-SOL-02:S1:S1_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S3", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v2", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans; file API required authentication and was not used", "raw_source_id": "S3", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 116951, "captured_sha256": "0a2c4c6ed54dd7e2e0ca6cc07aa19b18fa1439e506c2310db778094282d2cebc", "edition_id": "edition:keplinger-supplement-v2", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-02:S3:S3_SPAN_1", "V2-SOL-02:S3:S3_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S4", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S4", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 12692, "captured_sha256": "a2d811e0d24af89d9bd4b19cc9e8fbc474c67ef3827648550a6560eca2a5e332", "edition_id": "edition:drbench-iclr-2026", "http_status": 403, "media_type": "text/html", "redirect_chain": [], "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolved_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "retrieval_status": "inaccessible" }, "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-02:S4:S4_SPAN_1", "V2-SOL-02:S4:S4_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S5", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S5", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804v1", "resolved_url": "https://arxiv.org/html/2508.15804v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-SOL-02:S5:S5_SPAN_1", "V2-SOL-02:S5:S5_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S7", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S7", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635v1", "resolved_url": "https://arxiv.org/html/2605.06635v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-SOL-02:S7:S7_SPAN_1", "V2-SOL-02:S7:S7_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-03:s1", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5498, "captured_sha256": "d74c1891ae09a2b89121f2a60ac837c89761a8086759e45db214ca68c1905cea", "edition_id": "edition:keplinger-vor-2025", "http_status": 403, "media_type": "text/html; charset=UTF-8", "redirect_chain": [], "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolved_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "retrieval_status": "inaccessible" }, "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-03:s1:s1_span1" ] }, { "citation_occurrence_id": "V2-SOL-03:s3", "correction_ids": [], "edition_id": "edition:deeptrace-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 364455, "captured_sha256": "0c28edecaedd882584e985caac1e14f4a03dd59e939e17ae755a3c5e07ae42b0", "edition_id": "edition:deeptrace-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2509.04499", "resolved_url": "https://arxiv.org/html/2509.04499", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2509.04499", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-SOL-03:s3:s3_span1" ] }, { "citation_occurrence_id": "V2-SOL-03:s4", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173", "resolved_url": "https://arxiv.org/html/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-03:s4:s4_span1", "V2-SOL-03:s4:s4_span2", "V2-SOL-03:s4:s4_span3" ] }, { "citation_occurrence_id": "V2-SOL-04:S6", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v1", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans", "raw_source_id": "S6", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 115612, "captured_sha256": "da6308937337b2dfd3f97dc62f817e4a7ec362329ede0fce4c9bcdfd99ad93b1", "edition_id": "edition:keplinger-supplement-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "resolution_status": "unresolved", "run_id": "V2-SOL-04", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-04:S6:S6-SP1" ] }, { "citation_occurrence_id": "V2-TERRA-01:s1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 3678399, "captured_sha256": "8fbf30398f5e62f8839f0c9c8609bbb9e3cd0b57ae27d4bf33cb5db2007d1118", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolved_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-01:s1_deepresearchbench:s1_method", "V2-TERRA-01:s1_deepresearchbench:s1_table", "V2-TERRA-01:s1_deepresearchbench:s1_dates" ] }, { "citation_occurrence_id": "V2-TERRA-01:s2_reportbench", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s2_reportbench", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 729607, "captured_sha256": "90730ad75011d460305caf45a5be33b1b5eb4126f7e7efc1506f63beda1f91d3", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2508.15804", "resolved_url": "https://arxiv.org/pdf/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-01:s2_reportbench:s2_method", "V2-TERRA-01:s2_reportbench:s2_metric_table", "V2-TERRA-01:s2_reportbench:s2_collection" ] }, { "citation_occurrence_id": "V2-TERRA-01:s3_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 1566914, "captured_sha256": "1571435270d3e781ab8d5254913af00a7dcf0efcb99ec8e6d3db87a7b5c92f02", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2507.16280", "resolved_url": "https://arxiv.org/pdf/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-01:s3_researcherbench:s3_definition", "V2-TERRA-01:s3_researcherbench:s3_models_time", "V2-TERRA-01:s3_researcherbench:s3_table" ] }, { "citation_occurrence_id": "V2-TERRA-01:s4_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 1424589, "captured_sha256": "db5b7b7e3d9ce9fc6d6713b60f65f0d81b07090ded2a28be7f541d1c290230b5", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2605.06635", "resolved_url": "https://arxiv.org/pdf/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-01:s4_cited_not_verified:s4_definition_table", "V2-TERRA-01:s4_cited_not_verified:s4_depth" ] }, { "citation_occurrence_id": "V2-TERRA-01:s5_urlhealth", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s5_urlhealth", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 416944, "captured_sha256": "e8145a9f62f2a2bf0c5f3e2607160c56e010d02da965b5f65dbfa8f07ddde06c", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2604.03173", "resolved_url": "https://arxiv.org/pdf/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2604.03173", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-TERRA-01:s5_urlhealth:s5_scope", "V2-TERRA-01:s5_urlhealth:s5_table", "V2-TERRA-01:s5_urlhealth:s5_comparison" ] }, { "citation_occurrence_id": "V2-TERRA-02:s1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 3678399, "captured_sha256": "8fbf30398f5e62f8839f0c9c8609bbb9e3cd0b57ae27d4bf33cb5db2007d1118", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolved_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-02:s1_deepresearchbench:s1_fact_method", "V2-TERRA-02:s1_deepresearchbench:s1_fact_results", "V2-TERRA-02:s1_deepresearchbench:s1_fact_formula", "V2-TERRA-02:s1_deepresearchbench:s1_fact_judge_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s2_liveresearchbench", "correction_ids": [ "correction:liveresearchbench-license" ], "edition_id": "edition:liveresearchbench-arxiv-v5", "license": "CC BY-NC-SA 4.0", "license_treatment": "quote-minimal attributed spans; raw CC BY claim is corrected in review records", "raw_source_id": "s2_liveresearchbench", "raw_title": "LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild", "readback": { "captured_bytes": 13057891, "captured_sha256": "579b9728b76cfd242e9c94d9ff2985e196bbc72b5a741030e4f308ede04a4f69", "edition_id": "edition:liveresearchbench-arxiv-v5", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2510.14240v5", "resolved_url": "https://arxiv.org/pdf/2510.14240v5", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2510.14240v5", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:liveresearchbench", "span_occurrence_ids": [ "V2-TERRA-02:s2_liveresearchbench:s2_rubric_tree", "V2-TERRA-02:s2_liveresearchbench:s2_table7", "V2-TERRA-02:s2_liveresearchbench:s2_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s3_deeptrace", "correction_ids": [], "edition_id": "edition:deeptrace-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_deeptrace", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 2982839, "captured_sha256": "dea4981c1066d0240a005f603b3b14419c7e32beb5191b3df847ceb26af3d6b6", "edition_id": "edition:deeptrace-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-TERRA-02:s3_deeptrace:s3_definition", "V2-TERRA-02:s3_deeptrace:s3_table1", "V2-TERRA-02:s3_deeptrace:s3_corpus_and_retrieval", "V2-TERRA-02:s3_deeptrace:s3_judge_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s4_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s4_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 280402, "captured_sha256": "c080d7304274d70a39651f49297cc65910ca833d7ee6a3de3052370733e44ee3", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2507.16280", "resolved_url": "https://arxiv.org/html/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-02:s4_researcherbench:s4_method", "V2-TERRA-02:s4_researcherbench:s4_results" ] }, { "citation_occurrence_id": "V2-TERRA-03:S1", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635", "resolved_url": "https://arxiv.org/html/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-03:S1:S1-A", "V2-TERRA-03:S1:S1-B", "V2-TERRA-03:S1:S1-C", "V2-TERRA-03:S1:S1-D", "V2-TERRA-03:S1:S1-E", "V2-TERRA-03:S1:S1-F" ] }, { "citation_occurrence_id": "V2-TERRA-03:S2", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S2", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 352002, "captured_sha256": "9aa2894dbeaac30b23e7ffc8107a7f53b6e3855c8511838551de5bd7a422cc42", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2506.11763", "resolved_url": "https://arxiv.org/html/2506.11763", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2506.11763", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-03:S2:S2-A", "V2-TERRA-03:S2:S2-B", "V2-TERRA-03:S2:S2-C", "V2-TERRA-03:S2:S2-D" ] }, { "citation_occurrence_id": "V2-TERRA-03:S3", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S3", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804", "resolved_url": "https://arxiv.org/html/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-03:S3:S3-A", "V2-TERRA-03:S3:S3-B", "V2-TERRA-03:S3:S3-C", "V2-TERRA-03:S3:S3-D" ] }, { "citation_occurrence_id": "V2-TERRA-03:S4", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S4", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 111993, "captured_sha256": "abbaa9f22035230d13152a68aaeecd8e83e40867c9dae4acd1d547e36c866793", "edition_id": "edition:keplinger-vor-2025", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolved_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "retrieval_status": "retrieved" }, "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S4:S4-A", "V2-TERRA-03:S4:S4-B", "V2-TERRA-03:S4:S4-C" ] }, { "citation_occurrence_id": "V2-TERRA-03:S6", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S6", "raw_title": "PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5565, "captured_sha256": "a46109544fe4ff4504fab5e97abea3cb7172367aba54a308937386394f0ff046", "edition_id": "edition:keplinger-vor-2025", "http_status": 203, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolved_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "retrieval_status": "inaccessible" }, "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S6:S6-A" ] }, { "citation_occurrence_id": "V2-TERRA-04:S1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 352002, "captured_sha256": "9aa2894dbeaac30b23e7ffc8107a7f53b6e3855c8511838551de5bd7a422cc42", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2506.11763", "resolved_url": "https://arxiv.org/html/2506.11763", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2506.11763", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-04:S1_deepresearchbench:S1_method", "V2-TERRA-04:S1_deepresearchbench:S1_table1", "V2-TERRA-04:S1_deepresearchbench:S1_metric", "V2-TERRA-04:S1_deepresearchbench:S1_time" ] }, { "citation_occurrence_id": "V2-TERRA-04:S2_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S2_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 280402, "captured_sha256": "c080d7304274d70a39651f49297cc65910ca833d7ee6a3de3052370733e44ee3", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2507.16280", "resolved_url": "https://arxiv.org/html/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-04:S2_researcherbench:S2_method", "V2-TERRA-04:S2_researcherbench:S2_table2", "V2-TERRA-04:S2_researcherbench:S2_scope", "V2-TERRA-04:S2_researcherbench:S2_time" ] }, { "citation_occurrence_id": "V2-TERRA-04:S3_reportbench", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S3_reportbench", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804", "resolved_url": "https://arxiv.org/html/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-04:S3_reportbench:S3_method", "V2-TERRA-04:S3_reportbench:S3_table1", "V2-TERRA-04:S3_reportbench:S3_time", "V2-TERRA-04:S3_reportbench:S3_limitations" ] }, { "citation_occurrence_id": "V2-TERRA-04:S4_urlhealth", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S4_urlhealth", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173", "resolved_url": "https://arxiv.org/html/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-TERRA-04:S4_urlhealth:S4_table2", "V2-TERRA-04:S4_urlhealth:S4_comparison", "V2-TERRA-04:S4_urlhealth:S4_method", "V2-TERRA-04:S4_urlhealth:S4_limitations" ] }, { "citation_occurrence_id": "V2-TERRA-04:S5_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S5_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635", "resolved_url": "https://arxiv.org/html/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-04:S5_cited_not_verified:S5_method", "V2-TERRA-04:S5_cited_not_verified:S5_scope", "V2-TERRA-04:S5_cited_not_verified:S5_table1", "V2-TERRA-04:S5_cited_not_verified:S5_depth", "V2-TERRA-04:S5_cited_not_verified:S5_limitations" ] } ]
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:a399dcbcdf2310b4abfea59748d83a462942013aac90f34625abbff7c4f74a2e
- Span digest
- sha256:3187d8163b884569c89ff10ba36039614969488eaed19a9d52e9a3ff70148a20
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
04 The three inaccessible carrier occurrences remain visible and receive no credit.
Typed relation: undercutting
EM-0026 deterministic agent-citation evidence ledger
All inaccessible citation carriers
[ { "citation_occurrence_id": "V2-SOL-02:S4", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S4", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 12692, "captured_sha256": "a2d811e0d24af89d9bd4b19cc9e8fbc474c67ef3827648550a6560eca2a5e332", "edition_id": "edition:drbench-iclr-2026", "http_status": 403, "media_type": "text/html", "redirect_chain": [], "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolved_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "retrieval_status": "inaccessible" }, "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-02:S4:S4_SPAN_1", "V2-SOL-02:S4:S4_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-03:s1", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5498, "captured_sha256": "d74c1891ae09a2b89121f2a60ac837c89761a8086759e45db214ca68c1905cea", "edition_id": "edition:keplinger-vor-2025", "http_status": 403, "media_type": "text/html; charset=UTF-8", "redirect_chain": [], "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolved_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "retrieval_status": "inaccessible" }, "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-03:s1:s1_span1" ] }, { "citation_occurrence_id": "V2-TERRA-03:S6", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S6", "raw_title": "PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5565, "captured_sha256": "a46109544fe4ff4504fab5e97abea3cb7172367aba54a308937386394f0ff046", "edition_id": "edition:keplinger-vor-2025", "http_status": 203, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolved_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "retrieval_status": "inaccessible" }, "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S6:S6-A" ] } ]
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:fe4b10fe8b860e86a721cce5a2c2183c97c7fbdb3b69d7e5d9067a9c59783ae6
- Span digest
- sha256:98ee18b4c442784a4c37ac10bee772994b5e19978343aa8267197d4ab90028e6
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
05 The complete unsupported or force-raised set remains visible and receives no credit.
Typed relation: undercutting
EM-0026 deterministic agent-citation evidence ledger
Unsupported or force-raised claim occurrences
[ "V2-SOL-01:r3_deepresearch_bench_fact", "V2-SOL-01:r4_deeptrace_support", "V2-SOL-02:R3_deepresearch_bench_fact", "V2-SOL-02:R5_deeptrace_audit", "V2-SOL-03:r3_deepresearch_bench_fact", "V2-SOL-03:r4_deeptrace_support", "V2-SOL-03:r8_search_depth_ablation", "V2-SOL-03:r9_verifier_calibration", "V2-SOL-04:R1", "V2-SOL-04:R2", "V2-SOL-04:R3", "V2-SOL-04:R4", "V2-SOL-04:R5", "V2-SOL-04:R8", "V2-TERRA-01:r1_deepresearchbench_fact", "V2-TERRA-01:r4_cited_not_verified_source_attribution", "V2-TERRA-02:answer", "V2-TERRA-02:r1_deepresearchbench_fact", "V2-TERRA-02:r4_deeptrace", "V2-TERRA-03:R3" ]
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:087792c927d013e13f7eed7aafd137b4b759473b95866ccdd1fbf5ffb662e65f
- Span digest
- sha256:608e58805bcff58922f73ed2e0459b3cfcf1f8df9ee70e3d1238f68ac79712be
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
06 All nine independently rejected claims remain visible and receive no credit.
Typed relation: undercutting
EM-0026 deterministic agent-citation evidence ledger
Independently rejected claim occurrences
[ "V2-SOL-02:R5_deeptrace_audit", "V2-SOL-03:r3_deepresearch_bench_fact", "V2-SOL-03:r8_search_depth_ablation", "V2-SOL-03:r9_verifier_calibration", "V2-SOL-04:R1", "V2-SOL-04:R2", "V2-SOL-04:R3", "V2-SOL-04:R4", "V2-SOL-04:R8" ]
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:54fe13d731b79401537cd8dcc38d0c59a13c51a68916545018521cf600d772b7
- Span digest
- sha256:fc0a27fd59473c91e1a40670b6c9c07d037b5d6a56a2e977b7e7b8be5e240236
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
Verify every number
Complete count ledgers
Every displayed total is a view over these typed members; no total is maintained as marketing copy.
8 · Captured reports
V2-SOL-01V2-SOL-01 — completedV2-SOL-02V2-SOL-02 — completedV2-SOL-03V2-SOL-03 — completedV2-SOL-04V2-SOL-04 — completedV2-TERRA-01V2-TERRA-01 — completedV2-TERRA-02V2-TERRA-02 — completedV2-TERRA-03V2-TERRA-03 — completedV2-TERRA-04V2-TERRA-04 — completed
48 · Citation occurrences
V2-SOL-01:s1_keplingerV2-SOL-01:s1_keplinger — unresolvedV2-SOL-01:s1a_keplinger_dataV2-SOL-01:s1a_keplinger_data — unresolvedV2-SOL-01:s2_drbenchV2-SOL-01:s2_drbench — unresolvedV2-SOL-01:s2a_drbench_repoV2-SOL-01:s2a_drbench_repo — matched-exact-spanV2-SOL-01:s3_deeptraceV2-SOL-01:s3_deeptrace — unresolvedV2-SOL-01:s4_cited_not_verifiedV2-SOL-01:s4_cited_not_verified — unresolvedV2-SOL-01:s5_url_healthV2-SOL-01:s5_url_health — unresolvedV2-SOL-02:S1V2-SOL-02:S1 — unresolvedV2-SOL-02:S2V2-SOL-02:S2 — matched-exact-spanV2-SOL-02:S3V2-SOL-02:S3 — unresolvedV2-SOL-02:S4V2-SOL-02:S4 — unresolvedV2-SOL-02:S5V2-SOL-02:S5 — unresolvedV2-SOL-02:S6V2-SOL-02:S6 — matched-exact-spanV2-SOL-02:S7V2-SOL-02:S7 — unresolvedV2-SOL-03:s1V2-SOL-03:s1 — unresolvedV2-SOL-03:s2V2-SOL-03:s2 — matched-exact-spanV2-SOL-03:s3V2-SOL-03:s3 — unresolvedV2-SOL-03:s4V2-SOL-03:s4 — unresolvedV2-SOL-03:s5V2-SOL-03:s5 — matched-exact-spanV2-SOL-03:s6V2-SOL-03:s6 — matched-exact-spanV2-SOL-03:s7V2-SOL-03:s7 — matched-exact-spanV2-SOL-04:S1V2-SOL-04:S1 — matched-exact-spanV2-SOL-04:S2V2-SOL-04:S2 — matched-exact-spanV2-SOL-04:S3V2-SOL-04:S3 — matched-exact-spanV2-SOL-04:S4V2-SOL-04:S4 — matched-exact-spanV2-SOL-04:S5V2-SOL-04:S5 — matched-exact-spanV2-SOL-04:S6V2-SOL-04:S6 — unresolvedV2-SOL-04:S7V2-SOL-04:S7 — matched-exact-spanV2-TERRA-01:s1_deepresearchbenchV2-TERRA-01:s1_deepresearchbench — unresolvedV2-TERRA-01:s2_reportbenchV2-TERRA-01:s2_reportbench — unresolvedV2-TERRA-01:s3_researcherbenchV2-TERRA-01:s3_researcherbench — unresolvedV2-TERRA-01:s4_cited_not_verifiedV2-TERRA-01:s4_cited_not_verified — unresolvedV2-TERRA-01:s5_urlhealthV2-TERRA-01:s5_urlhealth — unresolvedV2-TERRA-02:s1_deepresearchbenchV2-TERRA-02:s1_deepresearchbench — unresolvedV2-TERRA-02:s2_liveresearchbenchV2-TERRA-02:s2_liveresearchbench — unresolvedV2-TERRA-02:s3_deeptraceV2-TERRA-02:s3_deeptrace — unresolvedV2-TERRA-02:s4_researcherbenchV2-TERRA-02:s4_researcherbench — unresolvedV2-TERRA-03:S1V2-TERRA-03:S1 — unresolvedV2-TERRA-03:S2V2-TERRA-03:S2 — unresolvedV2-TERRA-03:S3V2-TERRA-03:S3 — unresolvedV2-TERRA-03:S4V2-TERRA-03:S4 — unresolvedV2-TERRA-03:S5V2-TERRA-03:S5 — matched-exact-spanV2-TERRA-03:S6V2-TERRA-03:S6 — unresolvedV2-TERRA-04:S1_deepresearchbenchV2-TERRA-04:S1_deepresearchbench — unresolvedV2-TERRA-04:S2_researcherbenchV2-TERRA-04:S2_researcherbench — unresolvedV2-TERRA-04:S3_reportbenchV2-TERRA-04:S3_reportbench — unresolvedV2-TERRA-04:S4_urlhealthV2-TERRA-04:S4_urlhealth — unresolvedV2-TERRA-04:S5_cited_not_verifiedV2-TERRA-04:S5_cited_not_verified — unresolved
30 · Distinct cited URL strings
https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrievedhttps://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrievedhttps://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrievedhttps://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrievedhttps://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrievedhttps://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrievedhttps://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrievedhttps://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrievedhttps://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrievedhttps://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrievedhttps://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrievedhttps://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrievedhttps://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrievedhttps://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrievedhttps://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrievedhttps://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrievedhttps://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrievedhttps://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrievedhttps://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrievedhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrievedhttps://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrievedhttps://onlinelibrary.wiley.com/doi/10.1111/jdv.70035https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 — inaccessiblehttps://openreview.net/pdf?id=hQ0K2Hhq7Hhttps://openreview.net/pdf?id=hQ0K2Hhq7H — inaccessiblehttps://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrievedhttps://pubmed.ncbi.nlm.nih.gov/40904191/https://pubmed.ncbi.nlm.nih.gov/40904191/ — inaccessible
27 · Resolving URL roots
https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrievedhttps://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrievedhttps://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrievedhttps://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrievedhttps://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrievedhttps://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrievedhttps://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrievedhttps://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrievedhttps://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrievedhttps://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrievedhttps://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrievedhttps://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrievedhttps://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrievedhttps://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrievedhttps://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrievedhttps://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrievedhttps://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrievedhttps://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrievedhttps://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrievedhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrievedhttps://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrievedhttps://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
11 · Source works
work-citation-verifier-benchmark-2d5e94336bDo You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution — examined-source-workwork-cited-not-verified-3825b25622Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — examined-source-workwork-deepresearch-bench-paper-8ef36010d2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — examined-source-workwork-deepresearch-bench-repository-1e475c7631Ayanami0730/deep_research_bench — examined-source-workwork-deeptrace-650d66cfcbDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — examined-source-workwork-keplinger-dermatology-audit-639f9ea37aAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — examined-source-workwork-keplinger-supplement-1a03034fb3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — examined-source-workwork-liveresearchbench-61e7da557bLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — examined-source-workwork-reportbench-dca823c910ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — examined-source-workwork-researcherbench-82729ab60fResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — examined-source-workwork-url-health-f64342d493Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — examined-source-work
14 · Examined editions
edition-citation-verifier-arxiv-v1-acfd4abab7Quote-minimal projection of edition:citation-verifier-arxiv-v1 — examined-editionedition-cnv-arxiv-v1-f24001c50aQuote-minimal projection of edition:cnv-arxiv-v1 — examined-editionedition-deeptrace-arxiv-v1-70a6e720c1Quote-minimal projection of edition:deeptrace-arxiv-v1 — examined-editionedition-deeptrace-iclr-2026-b2b153eaa6Quote-minimal projection of edition:deeptrace-iclr-2026 — examined-editionedition-drbench-arxiv-v1-7afa5ade36Quote-minimal projection of edition:drbench-arxiv-v1 — examined-editionedition-drbench-iclr-2026-d00a3dcdb6Quote-minimal projection of edition:drbench-iclr-2026 — examined-editionedition-drbench-repo-main-469cce5-4f8964f3e3Quote-minimal projection of edition:drbench-repo-main-469cce5 — examined-editionedition-keplinger-supplement-v1-f6ddb4c04aQuote-minimal projection of edition:keplinger-supplement-v1 — examined-editionedition-keplinger-supplement-v2-01e44c09d8Quote-minimal projection of edition:keplinger-supplement-v2 — examined-editionedition-keplinger-vor-2025-10ba10f860Quote-minimal projection of edition:keplinger-vor-2025 — examined-editionedition-liveresearchbench-arxiv-v5-f5644763e5Quote-minimal projection of edition:liveresearchbench-arxiv-v5 — examined-editionedition-reportbench-arxiv-v1-25888dc44bQuote-minimal projection of edition:reportbench-arxiv-v1 — examined-editionedition-researcherbench-arxiv-v1-39a8b256d7Quote-minimal projection of edition:researcherbench-arxiv-v1 — examined-editionedition-url-health-arxiv-v1-85fb934b53Quote-minimal projection of edition:url-health-arxiv-v1 — examined-edition
72 · Accepted exact span roots
span-015bb68c904bf225e84c82b1acffb3cc38a4a087d7b0f79d-4c454092b1Evaluation methodology, lines 144-149 — matched-exact-spanspan-084f11e57cf779850ea87aaf30ba1277438ff9efcb533fd2-fe72b27cfbAppendix A.1 'Limitations' — matched-exact-spanspan-092745c7db6407d4d52927642741cde229fad05a5347bb18-7b04227f8fMain text, paragraph beginning 'For their high performances in generating reference lists' — matched-exact-spanspan-0abbab3d6269d58eb76d8e90eff1e4c27512cf46d9b8e5ff-54e5759a8brepository README, Overview — matched-exact-spanspan-0d4f357f2a5565bdaea01a1e537650958d71216868d1a2b0-9b0af35d76Table 1, lines 177-182 — matched-exact-spanspan-13cb0ddb087b82988b7445a9cfb02562f9c044ef4f244e8b-87c3b42c68Section 3.2, lines 144-146 — matched-exact-spanspan-196c54d6212fbc0c054dbf8e8467f2388c1f788bb93f4b1a-d460e927f6Table 1, Perplexity Deep Research citation-accuracy cell — matched-exact-spanspan-1db12b68cfeb4c5f5a96fa9cd2e2bd46e8dfb4f096f73154-cea47f0c65Limitations, lines 259-261 — matched-exact-spanspan-279db725380c133044944dccb379b75a396242219aa30437-569fc9e2fdpage 16, FACT judge validation — matched-exact-spanspan-2f311bfb451c6208101cb7583af37c338914fcdb10f5c75b-90e7764ce2Table 1, OpenAI Deep Research row; FACT columns — matched-exact-spanspan-3015a109e7898cb720e426e9a57c526277b648c75f1fb4b1-745d84ff44Section 4, line 186 — matched-exact-spanspan-35927851edd2e7e5e0f60d98498940f88304ba99bf1f85a0-2fe586436bpage 14, Table 3 — matched-exact-spanspan-39a9931f05ea356fe904ed15768fa08c66440b1fcb20f98e-191c24bbb2Table 1, ChatGPT all-correct subtotal — matched-exact-spanspan-417182b4c15ac550469e435a02aab1e91b395f72f186312d-e9e2c4a70cSection 4.3, line 212 — matched-exact-spanspan-46c392bc942cd88d525d74298e6bc6ebd9aeaec533ddfb7b-774f5b996emain text, paragraph beginning 'For their high performances'; Figure 1 — matched-exact-spanspan-55de6a6ca9414a252592e9f77544a9259533dd713ce1f12f-0826216ffcpp. 1-2, lines 64-95 — matched-exact-spanspan-59d0da64f738e1d8745f2bff3663eb89e273f93fe501afe3-ba69fc58ebSection 2 definition discussion and Section 3.3 — matched-exact-spanspan-5cd9ba4ca72da10cf51c0beddc6ea2765c74721bdd526a67-ed8f789629Limitations, lines 219-224 — matched-exact-spanspan-6885cf2524462022475b5829da081b2d35fcfe7463ca6ad0-3b1a122fafSection 1 contributions — matched-exact-spanspan-6c1d327e58ef001533ea1c11010e974436bbeafbde138efc-f6efd8d376Section 3.3, line 155 — matched-exact-spanspan-6e961498ed0f81a950f775956587f741e0eb9f11640fd3cb-afa5b59854Section 4.3, Tables 2–3 — matched-exact-spanspan-6f79c3c824ccf889d73e234335661e5fca68d90f8c78ee97-dbf7848a72dataset page, Description — matched-exact-spanspan-732c03e69c6e142a492275cee5d1e2d474c75422c9b84aaa-f7675998ccSection 3.3.1-3.3.3 — matched-exact-spanspan-74f6acafbb4c54d556a572460bf1b45bacd1e293f4aa16cc-5ee15c82c4Abstract, lines 47-48 — matched-exact-spanspan-76e9bd59b31c7ebd2ed7a2186e59753a79d3c730e0d7b4fb-9c98d99242Appendix D, Table 5 — matched-exact-spanspan-794bdf7e0595799afbb0d55f3dc522077aed9615c59f1e35-43f2ba944dAbstract, line 16 — matched-exact-spanspan-7a6046a756d6ec0315d330cf0817d062ea48498bdc9c90b3-c3a1e80029p. 5, lines 315-329 — matched-exact-spanspan-7d7331942f9d0520e18e96206e7718e5187788be4d48cef9-dfdbc1d019page 4, Section 3.2, Support Judgment — matched-exact-spanspan-886b50d94e3facabc487342cd2af5157cf956622ecfc3686-e7bdad0935Section 4.2.1, Citation Support Verification and Score Computation — matched-exact-spanspan-8c8d9a6dc003bf0298542b85df3f6f08db631e68e89f606f-92731379f2Table 1, ChatGPT subtotal — matched-exact-spanspan-8cce3d9ee1e1674ba1b93321dd9b99a55d834ca980451f8a-d2fceafe11Abstract, line 46 — matched-exact-spanspan-912c931063109f84ea88aea34891dd5e0f9147d2176af258-4d566ca076Abstract and Table 1 — matched-exact-spanspan-92eb7cf19745e24190dda84c77d590941134d09f06805c0f-88afe43377README FACT, lines 240-248 — matched-exact-spanspan-94ad583514aa7e04b8d5f7b8f9cdda6fdffe8ca6712f033d-63c1ba24ebDataset description — matched-exact-spanspan-9c7a4c186cc85f185aa293a7c5a46c08f9dbbc1a9e63c449-feff07a6e2Section 2.2 'Cited statements' — matched-exact-spanspan-9cb1700c08368feb234207045354cbf559d6ed80f2b0bae8-f898722eb5pp.4-7, Sections 3.1.1 and 3.2 — matched-exact-spanspan-9dd6eb46529e77a62748093528e52685fbd92d46a756aef7-96d75d9e2bSection 3.2, lines 163-171 — matched-exact-spanspan-9fe42cb48702d8d96ebbbc0309692214ece6ec495c447c06-80da541506Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment' — matched-exact-spanspan-a08468cb62f2feb21d9cbb95175ce00f54592bd83573cc91-03f65cfa45Official ICLR abstract — matched-exact-spanspan-a15ab02e9fcb23df02fb03d28967d4220ddf65ad380d2230-ddf85fa359page 4, source scraping — matched-exact-spanspan-a18e778fac3d49e81167f05e09fbc361e91af8c0c01b0fec-ca1b18e206page 3, Section 3.3 — matched-exact-spanspan-aa1534ce2e177153ab85e512deb8d064d14ec1eae16d8032-2299096a09p.14, Table 3 — matched-exact-spanspan-b0a39c963cdfdc5f9852cfc9f7e33e05d31232bd1d899885-adf9545222p. 4, lines 150-159 — matched-exact-spanspan-b38a2af5982b4c85bc205d2f533a23ed3a8f40a49b651a5f-27395a6291Main text, paragraph beginning 'To address the gap' (search-indexed primary full text) — matched-exact-spanspan-b68153666e4e12aee8dc887fef0a61fa5f44b5f680a0b7eb-769c2cc17dSection 5.1 Results — matched-exact-spanspan-b6a6ba48a2c4e352c8565b671a15a90400739f2aeefda8a6-e8232daf27Main text, following paragraph (search-indexed primary full text) — matched-exact-spanspan-b83996e3b5fdf878e04d6d41d0e7a1eebee9fdd9bbccf2ae-820a3eab86Appendix C, line 365 — matched-exact-spanspan-b903dc2284e1b4d80bb6ef948d79bbdbc7d7447e871281b8-6ef6234316body, paragraph immediately after Table 1 — matched-exact-spanspan-ba6b8a8b7d9121ebb05d594176ceed3d1e9dbc637876783a-d85147b6b0Settings and metrics, lines 132-139 — matched-exact-spanspan-bb16d5f06abe8641d9ea6694e549aa15e240dd88946aa47a-5946b6e6b7Abstract and Table 1 — matched-exact-spanspan-bb62928b5e0cb4e373b2f6bfeb579d8564d05da5588f8b15-771310f4efSection 3.3 'URL extraction and classification' — matched-exact-spanspan-c27aaca02bda06ed764be53351158fc862af9d4a556d3e82-104e7cda52page 5, Section 4.1 — matched-exact-spanspan-c3dfc98566c27561eb7267972e00e3e1ba6a399a3ff5d002-37f40d3260page 5, Section 3.3.3 — matched-exact-spanspan-c7dbd1d9e09980703ecbee8161594e4b426c1af63117b286-153d26f046Section 5 'Limitations' — matched-exact-spanspan-cd21c450c0d33df20a2b1a539d5725347db1b5bab7d1e986-2e3e0c3ac7Abstract, lines 45-49 — matched-exact-spanspan-d638c8b21f0e4a0f840880b6ffe0217ce1623758cc09fdcc-8585f6d39bSection 4.3, Tables 2-3 and following paragraph — matched-exact-spanspan-d880991ad784390029d49f84949d083c4789d67d45893d87-c6ef8a5c13p.33, Appendix E — matched-exact-spanspan-ddf4320a6d302c13d28b6225eb7e79ac15583349e1b449d4-f89bb46179Section 3.1 'Setttings' — matched-exact-spanspan-deada8bacb987a47f2a42e52fbfeb5c1beb3d250c0be79e2-0e91f5caccSection 4.4, line 217 — matched-exact-spanspan-e223640cb0ed12c87ca1d2406f3276f30a3b8d2017dd4dc1-3cfecd76bfTable 1, OpenAI Deep Research row and cited-statements Match Rate column — matched-exact-spanspan-e2c9f6220843a04665f5d1cd142ac15f04211ace5c579fa6-c5e0476df0Paper section 3.1.1, factual-support validation — matched-exact-spanspan-e4a9f30effc031388ed5c9993dc6f071903bec5e3f1998db-5d95269b76pp.17-18, Appendix E, equations 4-6 — matched-exact-spanspan-e54ee1457854a9a5e46f7c7c2e66ac6a5a69b963f8ee3450-8f875cfe94Abstract, line 16 — matched-exact-spanspan-e5bcf0241a1aeae7d9a92755b1bcf0adeeea93a8922f22f8-77a5ff4cc0README News, 2026-05-11 evaluator migration notice — matched-exact-spanspan-e6f447129d9a662fb202d4ac22e16822bacfdbcbcdb7afe3-f7abe2194cp.23, Appendix C, Citation Accuracy — matched-exact-spanspan-ea9cb2e1b0078c0857b789d61d0a0d26801cf66891864705-2f19633d83Section 5.1 and Appendix E, Table 4 — matched-exact-spanspan-eaf37bef7b8dc35f3e00b85d3e58ad872dc0ea604a062c8c-fa09300f85p.4, Section 3.2, lines 148-155 — matched-exact-spanspan-ef637b38f36195bea09205d532e8744daf3e257f93d97144-26007673a3README News, lines 176-188 — matched-exact-spanspan-f05b3189dc3fafd45b5bde643a91fd0616832928215a37a2-47f72e77e4Paragraph after Table 1 — matched-exact-spanspan-f35c9565f2fd047f6e7272dc608bbf6a63813a717bad5354-e6a07f05b6Appendix C, line 365 — matched-exact-spanspan-f96ad5daa81153a3689f52f5777fedfd70e8a87d44e7a1c9-3da72b8748Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3) — matched-exact-spanspan-f98f1274859ba98cc538e3be6d11d345a5382cfcb9516aa1-d60b765c48p.17, Appendix C, lines 704-708 — matched-exact-span
7 · Candidate warrant roots
relation-cnv-link-relevance-support-gap-165443d270Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.relation-cnv-search-depth-ablation-9c01321092Within the Cited but Not Verified harness, the 2-to-150-call ablation reduced reported Fact Check scores for two setups without establishing a general causal law about search depth.relation-deepresearch-bench-fact-a3f61d7a3aDeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.relation-deeptrace-support-variation-1a84ef7c6fDeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.relation-keplinger-metadata-versus-support-c23e105c8bIn one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.relation-reportbench-match-rate-6964e41a69ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.relation-url-health-resolution-0b26fb45c7The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.
0 · Independently confirmed warrant roots
4 · Pending warrant groups
relation-citation-verifier-calibration-5b40082faaThe captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-liveresearchbench-e1-e2-e3-5fb5d671f7The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-researcherbench-faithfulness-groundedness-c84dcaee54The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-url-health-correction-loop-eb8a6384c6The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
34 · Unresolved citation occurrences
V2-SOL-01:s1_keplingerAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-SOL-01:s1a_keplinger_dataSupplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-SOL-01:s2_drbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-01:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-SOL-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-SOL-01:s5_url_healthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-02:S1Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-02:S3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-02:S5ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-SOL-02:S7Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-SOL-03:s3DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-SOL-03:s4Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-04:S6Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-TERRA-01:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-01:s2_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-01:s3_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-TERRA-01:s5_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-TERRA-02:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-02:s2_liveresearchbenchLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — unresolvedV2-TERRA-02:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-TERRA-02:s4_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-03:S1Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-TERRA-03:S2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-03:S3ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-03:S4Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-04:S1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-04:S2_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-04:S3_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-04:S4_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-TERRA-04:S5_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
20 · Unsupported or force-raised claims
V2-SOL-01:r3_deepresearch_bench_factV2-SOL-01:r3_deepresearch_bench_fact — no-creditV2-SOL-01:r4_deeptrace_supportV2-SOL-01:r4_deeptrace_support — no-creditV2-SOL-02:R3_deepresearch_bench_factV2-SOL-02:R3_deepresearch_bench_fact — no-creditV2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — no-creditV2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — no-creditV2-SOL-03:r4_deeptrace_supportV2-SOL-03:r4_deeptrace_support — no-creditV2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — no-creditV2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — no-creditV2-SOL-04:R1V2-SOL-04:R1 — no-creditV2-SOL-04:R2V2-SOL-04:R2 — no-creditV2-SOL-04:R3V2-SOL-04:R3 — no-creditV2-SOL-04:R4V2-SOL-04:R4 — no-creditV2-SOL-04:R5V2-SOL-04:R5 — no-creditV2-SOL-04:R8V2-SOL-04:R8 — no-creditV2-TERRA-01:r1_deepresearchbench_factV2-TERRA-01:r1_deepresearchbench_fact — no-creditV2-TERRA-01:r4_cited_not_verified_source_attributionV2-TERRA-01:r4_cited_not_verified_source_attribution — no-creditV2-TERRA-02:answerV2-TERRA-02:answer — no-creditV2-TERRA-02:r1_deepresearchbench_factV2-TERRA-02:r1_deepresearchbench_fact — no-creditV2-TERRA-02:r4_deeptraceV2-TERRA-02:r4_deeptrace — no-creditV2-TERRA-03:R3V2-TERRA-03:R3 — no-credit
9 · Independently rejected claims
V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — independently-rejectedV2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — independently-rejectedV2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — independently-rejectedV2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — independently-rejectedV2-SOL-04:R1V2-SOL-04:R1 — independently-rejectedV2-SOL-04:R2V2-SOL-04:R2 — independently-rejectedV2-SOL-04:R3V2-SOL-04:R3 — independently-rejectedV2-SOL-04:R4V2-SOL-04:R4 — independently-rejectedV2-SOL-04:R8V2-SOL-04:R8 — independently-rejected
3 · Inaccessible carriers
V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
Five-line lexicon
The units are different on purpose
- Citation occurrence
- One report-level use of a citation, before repeated URLs are collapsed.
- URL root
- One distinct cited URL string; resolution does not prove claim support.
- Source work
- One logical paper, dataset, repository, or audit instrument across editions.
- Candidate warrant root
- One source-method-data proposition that survived bounded semantic review, while residual independence remains unresolved.
- No credit
- The item remains visible but does not support the stronger claim because its carrier, span, semantics, or lineage did not close.