How We Know · Case 002 · skeptical

When eight research agents agree, how many evidence roots are there?

What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?

Eight context-isolated public-by-design reports captured under one frozen prompt, plus the public source editions and exact spans they cited. This historical pilot does not estimate current or universal agent behavior.

skeptical finding

This pilot cannot estimate current agent citation reliability: thirty-four citation occurrences remain unresolved, twenty claims required correction or no credit, four warrant groups remain pending, and zero warrant roots are independently confirmed by the packet. Inspect the exact source span before relying on a polished cited answer.

What to do with that: Do not infer independent corroboration from eight agreeing reports: 34 citation occurrences remain unresolved, 20 claims lost credit, and no warrant root was independently confirmed.

The packet is a lineage audit, not a vendor ranking or a representative product test.

Lineage accounting

Agreement is not a vote count

Shared capture: Unknown provider and retrieval dependencies remain; all reports share the exact prompt and one bounded capture program, so run multiplicity gets zero automatic independence credit.

Warrant boundary: Unknown residual independence remains across task data, judge methods, retrieval, source, edition, span, derivation, and upstream-citation lineages.

Sentence x-ray

What the selected record actually supports

Open a sentence to inspect its exact work, edition, span, retrieval, digest, and license chain.

01 This report is a captured observation, not an independent evidence root.

Typed relation: dependence

EM-0026 deterministic agent-citation evidence ledger

Captured report V2-TERRA-04

{ "answer_bytes": 30189, "answer_path": "research/how-we-know/agent-citation-lineage/answers-v2/V2-TERRA-04.json", "answer_sha256": "8deffa91b6bb52d1de4ef4965d0087eca3d3a8a96d971a5e89753c1928cc4ef4", "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_model_profile": "gpt-5.6-terra", "retrieval_infrastructure": "unknown", "run_id": "V2-TERRA-04", "status": "completed", "trace_path": "research/how-we-know/agent-citation-lineage/traces-v2/V2-TERRA-04.json" }
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:58db8115527aba46c1706f0fc3999e811f0352f5e855ad09e56c8688118d94b8
Span digest
sha256:c9198194e5a51ebafd593b2d7e2386afb6ee76460f01bf2958f54df04981c528
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments

EM-0026 deterministic agent-citation evidence ledger

Shared capture lineage

{ "automatic_independence_credit": 0, "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_profiles": [ "gpt-5.6-sol", "gpt-5.6-terra" ], "retrieval_infrastructure": "unknown" }
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:e120064fee0bdabe15299cb865ff2f7df66473291a13b7a765a4a09d16223f5e
Span digest
sha256:92b558b23559a96a551de8acb2627b148688d21ede2ca3fb9966196c4675dab6
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
02 The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.

Typed relation: qualification

EM-0026 deterministic agent-citation evidence ledger

Pending warrant warrant:citation-verifier-calibration

warrant:citation-verifier-calibration
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:755999bc2f4f9858e910f7436ae3e011343fb8b295c7ac25ebe37124b9fa9a79
Span digest
sha256:da47a19a65e8c1a6816fe474c99558e32442ba728c9fe0046c1ee24d79ec33d4
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
03 The complete unresolved set remains visible and receives no warrant credit.

Typed relation: undercutting

EM-0026 deterministic agent-citation evidence ledger

All unresolved citation occurrences

[ { "citation_occurrence_id": "V2-SOL-01:s1_keplinger", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1_keplinger", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 111993, "captured_sha256": "abbaa9f22035230d13152a68aaeecd8e83e40867c9dae4acd1d547e36c866793", "edition_id": "edition:keplinger-vor-2025", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolved_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "retrieval_status": "retrieved" }, "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-01:s1_keplinger:sp1a", "V2-SOL-01:s1_keplinger:sp1b", "V2-SOL-01:s1_keplinger:sp1c" ] }, { "citation_occurrence_id": "V2-SOL-01:s1a_keplinger_data", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v2", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans; file API required authentication and was not used", "raw_source_id": "s1a_keplinger_data", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 116951, "captured_sha256": "0a2c4c6ed54dd7e2e0ca6cc07aa19b18fa1439e506c2310db778094282d2cebc", "edition_id": "edition:keplinger-supplement-v2", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-01:s1a_keplinger_data:sp1d" ] }, { "citation_occurrence_id": "V2-SOL-01:s2_drbench", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s2_drbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 2665857, "captured_sha256": "8f80ce247f7cc355bb6773f36037e01cc3ba2e2082c77381eaadf0d2b92b021c", "edition_id": "edition:drbench-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-01:s2_drbench:sp2a", "V2-SOL-01:s2_drbench:sp2b", "V2-SOL-01:s2_drbench:sp2c", "V2-SOL-01:s2_drbench:sp2d" ] }, { "citation_occurrence_id": "V2-SOL-01:s3_deeptrace", "correction_ids": [], "edition_id": "edition:deeptrace-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_deeptrace", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 2982839, "captured_sha256": "dea4981c1066d0240a005f603b3b14419c7e32beb5191b3df847ceb26af3d6b6", "edition_id": "edition:deeptrace-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-SOL-01:s3_deeptrace:sp3a", "V2-SOL-01:s3_deeptrace:sp3b", "V2-SOL-01:s3_deeptrace:sp3c", "V2-SOL-01:s3_deeptrace:sp3d" ] }, { "citation_occurrence_id": "V2-SOL-01:s4_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 43848, "captured_sha256": "d1b6d476b3e81460dbc8821757775a8fa589e0579b77167ef0f33d35cb819cb5", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/abs/2605.06635", "resolved_url": "https://arxiv.org/abs/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/abs/2605.06635", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-SOL-01:s4_cited_not_verified:sp4a", "V2-SOL-01:s4_cited_not_verified:sp4b", "V2-SOL-01:s4_cited_not_verified:sp4c" ] }, { "citation_occurrence_id": "V2-SOL-01:s5_url_health", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s5_url_health", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 42408, "captured_sha256": "3976893e82f91cc9e7826c0d2d08c6caab40778134dbab09ebbfe1f6f63f5397", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/abs/2604.03173", "resolved_url": "https://arxiv.org/abs/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/abs/2604.03173", "resolution_status": "unresolved", "run_id": "V2-SOL-01", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-01:s5_url_health:sp5a", "V2-SOL-01:s5_url_health:sp5b", "V2-SOL-01:s5_url_health:sp5c" ] }, { "citation_occurrence_id": "V2-SOL-02:S1", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173v1", "resolved_url": "https://arxiv.org/html/2604.03173v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-02:S1:S1_SPAN_1", "V2-SOL-02:S1:S1_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S3", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v2", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans; file API required authentication and was not used", "raw_source_id": "S3", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 116951, "captured_sha256": "0a2c4c6ed54dd7e2e0ca6cc07aa19b18fa1439e506c2310db778094282d2cebc", "edition_id": "edition:keplinger-supplement-v2", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/2", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-02:S3:S3_SPAN_1", "V2-SOL-02:S3:S3_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S4", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S4", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 12692, "captured_sha256": "a2d811e0d24af89d9bd4b19cc9e8fbc474c67ef3827648550a6560eca2a5e332", "edition_id": "edition:drbench-iclr-2026", "http_status": 403, "media_type": "text/html", "redirect_chain": [], "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolved_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "retrieval_status": "inaccessible" }, "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-02:S4:S4_SPAN_1", "V2-SOL-02:S4:S4_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S5", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S5", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804v1", "resolved_url": "https://arxiv.org/html/2508.15804v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-SOL-02:S5:S5_SPAN_1", "V2-SOL-02:S5:S5_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-02:S7", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S7", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635v1", "resolved_url": "https://arxiv.org/html/2605.06635v1", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635v1", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-SOL-02:S7:S7_SPAN_1", "V2-SOL-02:S7:S7_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-03:s1", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5498, "captured_sha256": "d74c1891ae09a2b89121f2a60ac837c89761a8086759e45db214ca68c1905cea", "edition_id": "edition:keplinger-vor-2025", "http_status": 403, "media_type": "text/html; charset=UTF-8", "redirect_chain": [], "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolved_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "retrieval_status": "inaccessible" }, "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-03:s1:s1_span1" ] }, { "citation_occurrence_id": "V2-SOL-03:s3", "correction_ids": [], "edition_id": "edition:deeptrace-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 364455, "captured_sha256": "0c28edecaedd882584e985caac1e14f4a03dd59e939e17ae755a3c5e07ae42b0", "edition_id": "edition:deeptrace-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2509.04499", "resolved_url": "https://arxiv.org/html/2509.04499", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2509.04499", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-SOL-03:s3:s3_span1" ] }, { "citation_occurrence_id": "V2-SOL-03:s4", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173", "resolved_url": "https://arxiv.org/html/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-SOL-03:s4:s4_span1", "V2-SOL-03:s4:s4_span2", "V2-SOL-03:s4:s4_span3" ] }, { "citation_occurrence_id": "V2-SOL-04:S6", "correction_ids": [ "correction:mendeley-file-readback" ], "edition_id": "edition:keplinger-supplement-v1", "license": "CC BY 4.0", "license_treatment": "metadata and quote-minimal landing-page spans", "raw_source_id": "S6", "raw_title": "Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype", "readback": { "captured_bytes": 115612, "captured_sha256": "da6308937337b2dfd3f97dc62f817e4a7ec362329ede0fce4c9bcdfd99ad93b1", "edition_id": "edition:keplinger-supplement-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "resolved_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "retrieval_status": "retrieved" }, "requested_url": "https://data.mendeley.com/datasets/3s73z9zf3c/1", "resolution_status": "unresolved", "run_id": "V2-SOL-04", "source_work_id": "work:keplinger-supplement", "span_occurrence_ids": [ "V2-SOL-04:S6:S6-SP1" ] }, { "citation_occurrence_id": "V2-TERRA-01:s1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 3678399, "captured_sha256": "8fbf30398f5e62f8839f0c9c8609bbb9e3cd0b57ae27d4bf33cb5db2007d1118", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolved_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-01:s1_deepresearchbench:s1_method", "V2-TERRA-01:s1_deepresearchbench:s1_table", "V2-TERRA-01:s1_deepresearchbench:s1_dates" ] }, { "citation_occurrence_id": "V2-TERRA-01:s2_reportbench", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s2_reportbench", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 729607, "captured_sha256": "90730ad75011d460305caf45a5be33b1b5eb4126f7e7efc1506f63beda1f91d3", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2508.15804", "resolved_url": "https://arxiv.org/pdf/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-01:s2_reportbench:s2_method", "V2-TERRA-01:s2_reportbench:s2_metric_table", "V2-TERRA-01:s2_reportbench:s2_collection" ] }, { "citation_occurrence_id": "V2-TERRA-01:s3_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 1566914, "captured_sha256": "1571435270d3e781ab8d5254913af00a7dcf0efcb99ec8e6d3db87a7b5c92f02", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2507.16280", "resolved_url": "https://arxiv.org/pdf/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-01:s3_researcherbench:s3_definition", "V2-TERRA-01:s3_researcherbench:s3_models_time", "V2-TERRA-01:s3_researcherbench:s3_table" ] }, { "citation_occurrence_id": "V2-TERRA-01:s4_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s4_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 1424589, "captured_sha256": "db5b7b7e3d9ce9fc6d6713b60f65f0d81b07090ded2a28be7f541d1c290230b5", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2605.06635", "resolved_url": "https://arxiv.org/pdf/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-01:s4_cited_not_verified:s4_definition_table", "V2-TERRA-01:s4_cited_not_verified:s4_depth" ] }, { "citation_occurrence_id": "V2-TERRA-01:s5_urlhealth", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s5_urlhealth", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 416944, "captured_sha256": "e8145a9f62f2a2bf0c5f3e2607160c56e010d02da965b5f65dbfa8f07ddde06c", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2604.03173", "resolved_url": "https://arxiv.org/pdf/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2604.03173", "resolution_status": "unresolved", "run_id": "V2-TERRA-01", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-TERRA-01:s5_urlhealth:s5_scope", "V2-TERRA-01:s5_urlhealth:s5_table", "V2-TERRA-01:s5_urlhealth:s5_comparison" ] }, { "citation_occurrence_id": "V2-TERRA-02:s1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "s1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 3678399, "captured_sha256": "8fbf30398f5e62f8839f0c9c8609bbb9e3cd0b57ae27d4bf33cb5db2007d1118", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolved_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-02:s1_deepresearchbench:s1_fact_method", "V2-TERRA-02:s1_deepresearchbench:s1_fact_results", "V2-TERRA-02:s1_deepresearchbench:s1_fact_formula", "V2-TERRA-02:s1_deepresearchbench:s1_fact_judge_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s2_liveresearchbench", "correction_ids": [ "correction:liveresearchbench-license" ], "edition_id": "edition:liveresearchbench-arxiv-v5", "license": "CC BY-NC-SA 4.0", "license_treatment": "quote-minimal attributed spans; raw CC BY claim is corrected in review records", "raw_source_id": "s2_liveresearchbench", "raw_title": "LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild", "readback": { "captured_bytes": 13057891, "captured_sha256": "579b9728b76cfd242e9c94d9ff2985e196bbc72b5a741030e4f308ede04a4f69", "edition_id": "edition:liveresearchbench-arxiv-v5", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://arxiv.org/pdf/2510.14240v5", "resolved_url": "https://arxiv.org/pdf/2510.14240v5", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/pdf/2510.14240v5", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:liveresearchbench", "span_occurrence_ids": [ "V2-TERRA-02:s2_liveresearchbench:s2_rubric_tree", "V2-TERRA-02:s2_liveresearchbench:s2_table7", "V2-TERRA-02:s2_liveresearchbench:s2_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s3_deeptrace", "correction_ids": [], "edition_id": "edition:deeptrace-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s3_deeptrace", "raw_title": "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence", "readback": { "captured_bytes": 2982839, "captured_sha256": "dea4981c1066d0240a005f603b3b14419c7e32beb5191b3df847ceb26af3d6b6", "edition_id": "edition:deeptrace-iclr-2026", "http_status": 200, "media_type": "application/pdf", "redirect_chain": [], "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolved_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "retrieval_status": "retrieved" }, "requested_url": "https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:deeptrace", "span_occurrence_ids": [ "V2-TERRA-02:s3_deeptrace:s3_definition", "V2-TERRA-02:s3_deeptrace:s3_table1", "V2-TERRA-02:s3_deeptrace:s3_corpus_and_retrieval", "V2-TERRA-02:s3_deeptrace:s3_judge_validation" ] }, { "citation_occurrence_id": "V2-TERRA-02:s4_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "s4_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 280402, "captured_sha256": "c080d7304274d70a39651f49297cc65910ca833d7ee6a3de3052370733e44ee3", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2507.16280", "resolved_url": "https://arxiv.org/html/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-02", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-02:s4_researcherbench:s4_method", "V2-TERRA-02:s4_researcherbench:s4_results" ] }, { "citation_occurrence_id": "V2-TERRA-03:S1", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635", "resolved_url": "https://arxiv.org/html/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-03:S1:S1-A", "V2-TERRA-03:S1:S1-B", "V2-TERRA-03:S1:S1-C", "V2-TERRA-03:S1:S1-D", "V2-TERRA-03:S1:S1-E", "V2-TERRA-03:S1:S1-F" ] }, { "citation_occurrence_id": "V2-TERRA-03:S2", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S2", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 352002, "captured_sha256": "9aa2894dbeaac30b23e7ffc8107a7f53b6e3855c8511838551de5bd7a422cc42", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2506.11763", "resolved_url": "https://arxiv.org/html/2506.11763", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2506.11763", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-03:S2:S2-A", "V2-TERRA-03:S2:S2-B", "V2-TERRA-03:S2:S2-C", "V2-TERRA-03:S2:S2-D" ] }, { "citation_occurrence_id": "V2-TERRA-03:S3", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S3", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804", "resolved_url": "https://arxiv.org/html/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-03:S3:S3-A", "V2-TERRA-03:S3:S3-B", "V2-TERRA-03:S3:S3-C", "V2-TERRA-03:S3:S3-D" ] }, { "citation_occurrence_id": "V2-TERRA-03:S4", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S4", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 111993, "captured_sha256": "abbaa9f22035230d13152a68aaeecd8e83e40867c9dae4acd1d547e36c866793", "edition_id": "edition:keplinger-vor-2025", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolved_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "retrieval_status": "retrieved" }, "requested_url": "https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S4:S4-A", "V2-TERRA-03:S4:S4-B", "V2-TERRA-03:S4:S4-C" ] }, { "citation_occurrence_id": "V2-TERRA-03:S6", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S6", "raw_title": "PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5565, "captured_sha256": "a46109544fe4ff4504fab5e97abea3cb7172367aba54a308937386394f0ff046", "edition_id": "edition:keplinger-vor-2025", "http_status": 203, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolved_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "retrieval_status": "inaccessible" }, "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S6:S6-A" ] }, { "citation_occurrence_id": "V2-TERRA-04:S1_deepresearchbench", "correction_ids": [], "edition_id": "edition:drbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S1_deepresearchbench", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 352002, "captured_sha256": "9aa2894dbeaac30b23e7ffc8107a7f53b6e3855c8511838551de5bd7a422cc42", "edition_id": "edition:drbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2506.11763", "resolved_url": "https://arxiv.org/html/2506.11763", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2506.11763", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-TERRA-04:S1_deepresearchbench:S1_method", "V2-TERRA-04:S1_deepresearchbench:S1_table1", "V2-TERRA-04:S1_deepresearchbench:S1_metric", "V2-TERRA-04:S1_deepresearchbench:S1_time" ] }, { "citation_occurrence_id": "V2-TERRA-04:S2_researcherbench", "correction_ids": [], "edition_id": "edition:researcherbench-arxiv-v1", "license": "arXiv non-exclusive distribution license", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S2_researcherbench", "raw_title": "ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry", "readback": { "captured_bytes": 280402, "captured_sha256": "c080d7304274d70a39651f49297cc65910ca833d7ee6a3de3052370733e44ee3", "edition_id": "edition:researcherbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2507.16280", "resolved_url": "https://arxiv.org/html/2507.16280", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2507.16280", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:researcherbench", "span_occurrence_ids": [ "V2-TERRA-04:S2_researcherbench:S2_method", "V2-TERRA-04:S2_researcherbench:S2_table2", "V2-TERRA-04:S2_researcherbench:S2_scope", "V2-TERRA-04:S2_researcherbench:S2_time" ] }, { "citation_occurrence_id": "V2-TERRA-04:S3_reportbench", "correction_ids": [], "edition_id": "edition:reportbench-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S3_reportbench", "raw_title": "ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks", "readback": { "captured_bytes": 157386, "captured_sha256": "055e568189a402dd510c7c84be60f59465d9a001808936bc25cb0026ac25d267", "edition_id": "edition:reportbench-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2508.15804", "resolved_url": "https://arxiv.org/html/2508.15804", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2508.15804", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:reportbench", "span_occurrence_ids": [ "V2-TERRA-04:S3_reportbench:S3_method", "V2-TERRA-04:S3_reportbench:S3_table1", "V2-TERRA-04:S3_reportbench:S3_time", "V2-TERRA-04:S3_reportbench:S3_limitations" ] }, { "citation_occurrence_id": "V2-TERRA-04:S4_urlhealth", "correction_ids": [], "edition_id": "edition:url-health-arxiv-v1", "license": "CC0 1.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S4_urlhealth", "raw_title": "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents", "readback": { "captured_bytes": 302498, "captured_sha256": "332e5b5cb4b0ee7065b1bbc30436dfdfeafca8130fb87e32b0020026d8bbafe1", "edition_id": "edition:url-health-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2604.03173", "resolved_url": "https://arxiv.org/html/2604.03173", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2604.03173", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:url-health", "span_occurrence_ids": [ "V2-TERRA-04:S4_urlhealth:S4_table2", "V2-TERRA-04:S4_urlhealth:S4_comparison", "V2-TERRA-04:S4_urlhealth:S4_method", "V2-TERRA-04:S4_urlhealth:S4_limitations" ] }, { "citation_occurrence_id": "V2-TERRA-04:S5_cited_not_verified", "correction_ids": [], "edition_id": "edition:cnv-arxiv-v1", "license": "CC BY 4.0", "license_treatment": "quote-minimal attributed spans", "raw_source_id": "S5_cited_not_verified", "raw_title": "Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents", "readback": { "captured_bytes": 138385, "captured_sha256": "7c5e3c33f3122b07d176e6679cd6762a95a6babaf4a353686f019f905b5526ca", "edition_id": "edition:cnv-arxiv-v1", "http_status": 200, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://arxiv.org/html/2605.06635", "resolved_url": "https://arxiv.org/html/2605.06635", "retrieval_status": "retrieved" }, "requested_url": "https://arxiv.org/html/2605.06635", "resolution_status": "unresolved", "run_id": "V2-TERRA-04", "source_work_id": "work:cited-not-verified", "span_occurrence_ids": [ "V2-TERRA-04:S5_cited_not_verified:S5_method", "V2-TERRA-04:S5_cited_not_verified:S5_scope", "V2-TERRA-04:S5_cited_not_verified:S5_table1", "V2-TERRA-04:S5_cited_not_verified:S5_depth", "V2-TERRA-04:S5_cited_not_verified:S5_limitations" ] } ]
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:a399dcbcdf2310b4abfea59748d83a462942013aac90f34625abbff7c4f74a2e
Span digest
sha256:3187d8163b884569c89ff10ba36039614969488eaed19a9d52e9a3ff70148a20
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
04 The three inaccessible carrier occurrences remain visible and receive no credit.

Typed relation: undercutting

EM-0026 deterministic agent-citation evidence ledger

All inaccessible citation carriers

[ { "citation_occurrence_id": "V2-SOL-02:S4", "correction_ids": [], "edition_id": "edition:drbench-iclr-2026", "license": "unknown", "license_treatment": "metadata, digest, and quote-minimal attributed spans only", "raw_source_id": "S4", "raw_title": "DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents", "readback": { "captured_bytes": 12692, "captured_sha256": "a2d811e0d24af89d9bd4b19cc9e8fbc474c67ef3827648550a6560eca2a5e332", "edition_id": "edition:drbench-iclr-2026", "http_status": 403, "media_type": "text/html", "redirect_chain": [], "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolved_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "retrieval_status": "inaccessible" }, "requested_url": "https://openreview.net/pdf?id=hQ0K2Hhq7H", "resolution_status": "unresolved", "run_id": "V2-SOL-02", "source_work_id": "work:deepresearch-bench-paper", "span_occurrence_ids": [ "V2-SOL-02:S4:S4_SPAN_1", "V2-SOL-02:S4:S4_SPAN_2" ] }, { "citation_occurrence_id": "V2-SOL-03:s1", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "s1", "raw_title": "Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5498, "captured_sha256": "d74c1891ae09a2b89121f2a60ac837c89761a8086759e45db214ca68c1905cea", "edition_id": "edition:keplinger-vor-2025", "http_status": 403, "media_type": "text/html; charset=UTF-8", "redirect_chain": [], "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolved_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "retrieval_status": "inaccessible" }, "requested_url": "https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035", "resolution_status": "unresolved", "run_id": "V2-SOL-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-SOL-03:s1:s1_span1" ] }, { "citation_occurrence_id": "V2-TERRA-03:S6", "correction_ids": [], "edition_id": "edition:keplinger-vor-2025", "license": "CC BY-NC 4.0", "license_treatment": "quote-minimal attributed spans; no full-text redistribution", "raw_source_id": "S6", "raw_title": "PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype", "readback": { "captured_bytes": 5565, "captured_sha256": "a46109544fe4ff4504fab5e97abea3cb7172367aba54a308937386394f0ff046", "edition_id": "edition:keplinger-vor-2025", "http_status": 203, "media_type": "text/html; charset=utf-8", "redirect_chain": [], "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolved_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "retrieval_status": "inaccessible" }, "requested_url": "https://pubmed.ncbi.nlm.nih.gov/40904191/", "resolution_status": "unresolved", "run_id": "V2-TERRA-03", "source_work_id": "work:keplinger-dermatology-audit", "span_occurrence_ids": [ "V2-TERRA-03:S6:S6-A" ] } ]
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:fe4b10fe8b860e86a721cce5a2c2183c97c7fbdb3b69d7e5d9067a9c59783ae6
Span digest
sha256:98ee18b4c442784a4c37ac10bee772994b5e19978343aa8267197d4ab90028e6
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
05 The complete unsupported or force-raised set remains visible and receives no credit.

Typed relation: undercutting

EM-0026 deterministic agent-citation evidence ledger

Unsupported or force-raised claim occurrences

[ "V2-SOL-01:r3_deepresearch_bench_fact", "V2-SOL-01:r4_deeptrace_support", "V2-SOL-02:R3_deepresearch_bench_fact", "V2-SOL-02:R5_deeptrace_audit", "V2-SOL-03:r3_deepresearch_bench_fact", "V2-SOL-03:r4_deeptrace_support", "V2-SOL-03:r8_search_depth_ablation", "V2-SOL-03:r9_verifier_calibration", "V2-SOL-04:R1", "V2-SOL-04:R2", "V2-SOL-04:R3", "V2-SOL-04:R4", "V2-SOL-04:R5", "V2-SOL-04:R8", "V2-TERRA-01:r1_deepresearchbench_fact", "V2-TERRA-01:r4_cited_not_verified_source_attribution", "V2-TERRA-02:answer", "V2-TERRA-02:r1_deepresearchbench_fact", "V2-TERRA-02:r4_deeptrace", "V2-TERRA-03:R3" ]
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:087792c927d013e13f7eed7aafd137b4b759473b95866ccdd1fbf5ffb662e65f
Span digest
sha256:608e58805bcff58922f73ed2e0459b3cfcf1f8df9ee70e3d1238f68ac79712be
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
06 All nine independently rejected claims remain visible and receive no credit.

Typed relation: undercutting

EM-0026 deterministic agent-citation evidence ledger

Independently rejected claim occurrences

[ "V2-SOL-02:R5_deeptrace_audit", "V2-SOL-03:r3_deepresearch_bench_fact", "V2-SOL-03:r8_search_depth_ablation", "V2-SOL-03:r9_verifier_calibration", "V2-SOL-04:R1", "V2-SOL-04:R2", "V2-SOL-04:R3", "V2-SOL-04:R4", "V2-SOL-04:R8" ]
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:54fe13d731b79401537cd8dcc38d0c59a13c51a68916545018521cf600d772b7
Span digest
sha256:fc0a27fd59473c91e1a40670b6c9c07d037b5d6a56a2e977b7e7b8be5e240236
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments

Verify every number

Complete count ledgers

Every displayed total is a view over these typed members; no total is maintained as marketing copy.

8 · Captured reports
  1. V2-SOL-01V2-SOL-01 — completed
  2. V2-SOL-02V2-SOL-02 — completed
  3. V2-SOL-03V2-SOL-03 — completed
  4. V2-SOL-04V2-SOL-04 — completed
  5. V2-TERRA-01V2-TERRA-01 — completed
  6. V2-TERRA-02V2-TERRA-02 — completed
  7. V2-TERRA-03V2-TERRA-03 — completed
  8. V2-TERRA-04V2-TERRA-04 — completed
48 · Citation occurrences
  1. V2-SOL-01:s1_keplingerV2-SOL-01:s1_keplinger — unresolved
  2. V2-SOL-01:s1a_keplinger_dataV2-SOL-01:s1a_keplinger_data — unresolved
  3. V2-SOL-01:s2_drbenchV2-SOL-01:s2_drbench — unresolved
  4. V2-SOL-01:s2a_drbench_repoV2-SOL-01:s2a_drbench_repo — matched-exact-span
  5. V2-SOL-01:s3_deeptraceV2-SOL-01:s3_deeptrace — unresolved
  6. V2-SOL-01:s4_cited_not_verifiedV2-SOL-01:s4_cited_not_verified — unresolved
  7. V2-SOL-01:s5_url_healthV2-SOL-01:s5_url_health — unresolved
  8. V2-SOL-02:S1V2-SOL-02:S1 — unresolved
  9. V2-SOL-02:S2V2-SOL-02:S2 — matched-exact-span
  10. V2-SOL-02:S3V2-SOL-02:S3 — unresolved
  11. V2-SOL-02:S4V2-SOL-02:S4 — unresolved
  12. V2-SOL-02:S5V2-SOL-02:S5 — unresolved
  13. V2-SOL-02:S6V2-SOL-02:S6 — matched-exact-span
  14. V2-SOL-02:S7V2-SOL-02:S7 — unresolved
  15. V2-SOL-03:s1V2-SOL-03:s1 — unresolved
  16. V2-SOL-03:s2V2-SOL-03:s2 — matched-exact-span
  17. V2-SOL-03:s3V2-SOL-03:s3 — unresolved
  18. V2-SOL-03:s4V2-SOL-03:s4 — unresolved
  19. V2-SOL-03:s5V2-SOL-03:s5 — matched-exact-span
  20. V2-SOL-03:s6V2-SOL-03:s6 — matched-exact-span
  21. V2-SOL-03:s7V2-SOL-03:s7 — matched-exact-span
  22. V2-SOL-04:S1V2-SOL-04:S1 — matched-exact-span
  23. V2-SOL-04:S2V2-SOL-04:S2 — matched-exact-span
  24. V2-SOL-04:S3V2-SOL-04:S3 — matched-exact-span
  25. V2-SOL-04:S4V2-SOL-04:S4 — matched-exact-span
  26. V2-SOL-04:S5V2-SOL-04:S5 — matched-exact-span
  27. V2-SOL-04:S6V2-SOL-04:S6 — unresolved
  28. V2-SOL-04:S7V2-SOL-04:S7 — matched-exact-span
  29. V2-TERRA-01:s1_deepresearchbenchV2-TERRA-01:s1_deepresearchbench — unresolved
  30. V2-TERRA-01:s2_reportbenchV2-TERRA-01:s2_reportbench — unresolved
  31. V2-TERRA-01:s3_researcherbenchV2-TERRA-01:s3_researcherbench — unresolved
  32. V2-TERRA-01:s4_cited_not_verifiedV2-TERRA-01:s4_cited_not_verified — unresolved
  33. V2-TERRA-01:s5_urlhealthV2-TERRA-01:s5_urlhealth — unresolved
  34. V2-TERRA-02:s1_deepresearchbenchV2-TERRA-02:s1_deepresearchbench — unresolved
  35. V2-TERRA-02:s2_liveresearchbenchV2-TERRA-02:s2_liveresearchbench — unresolved
  36. V2-TERRA-02:s3_deeptraceV2-TERRA-02:s3_deeptrace — unresolved
  37. V2-TERRA-02:s4_researcherbenchV2-TERRA-02:s4_researcherbench — unresolved
  38. V2-TERRA-03:S1V2-TERRA-03:S1 — unresolved
  39. V2-TERRA-03:S2V2-TERRA-03:S2 — unresolved
  40. V2-TERRA-03:S3V2-TERRA-03:S3 — unresolved
  41. V2-TERRA-03:S4V2-TERRA-03:S4 — unresolved
  42. V2-TERRA-03:S5V2-TERRA-03:S5 — matched-exact-span
  43. V2-TERRA-03:S6V2-TERRA-03:S6 — unresolved
  44. V2-TERRA-04:S1_deepresearchbenchV2-TERRA-04:S1_deepresearchbench — unresolved
  45. V2-TERRA-04:S2_researcherbenchV2-TERRA-04:S2_researcherbench — unresolved
  46. V2-TERRA-04:S3_reportbenchV2-TERRA-04:S3_reportbench — unresolved
  47. V2-TERRA-04:S4_urlhealthV2-TERRA-04:S4_urlhealth — unresolved
  48. V2-TERRA-04:S5_cited_not_verifiedV2-TERRA-04:S5_cited_not_verified — unresolved
30 · Distinct cited URL strings
  1. https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrieved
  2. https://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrieved
  3. https://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrieved
  4. https://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrieved
  5. https://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrieved
  6. https://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrieved
  7. https://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrieved
  8. https://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrieved
  9. https://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrieved
  10. https://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrieved
  11. https://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrieved
  12. https://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrieved
  13. https://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrieved
  14. https://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrieved
  15. https://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrieved
  16. https://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrieved
  17. https://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrieved
  18. https://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrieved
  19. https://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrieved
  20. https://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
  21. https://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
  22. https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
  23. https://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrieved
  24. https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 — inaccessible
  25. https://openreview.net/pdf?id=hQ0K2Hhq7Hhttps://openreview.net/pdf?id=hQ0K2Hhq7H — inaccessible
  26. https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
  27. https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
  28. https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
  29. https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
  30. https://pubmed.ncbi.nlm.nih.gov/40904191/https://pubmed.ncbi.nlm.nih.gov/40904191/ — inaccessible
27 · Resolving URL roots
  1. https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrieved
  2. https://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrieved
  3. https://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrieved
  4. https://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrieved
  5. https://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrieved
  6. https://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrieved
  7. https://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrieved
  8. https://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrieved
  9. https://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrieved
  10. https://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrieved
  11. https://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrieved
  12. https://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrieved
  13. https://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrieved
  14. https://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrieved
  15. https://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrieved
  16. https://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrieved
  17. https://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrieved
  18. https://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrieved
  19. https://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrieved
  20. https://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
  21. https://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
  22. https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
  23. https://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrieved
  24. https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
  25. https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
  26. https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
  27. https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
11 · Source works
  1. work-citation-verifier-benchmark-2d5e94336bDo You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution — examined-source-work
  2. work-cited-not-verified-3825b25622Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — examined-source-work
  3. work-deepresearch-bench-paper-8ef36010d2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — examined-source-work
  4. work-deepresearch-bench-repository-1e475c7631Ayanami0730/deep_research_bench — examined-source-work
  5. work-deeptrace-650d66cfcbDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — examined-source-work
  6. work-keplinger-dermatology-audit-639f9ea37aAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — examined-source-work
  7. work-keplinger-supplement-1a03034fb3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — examined-source-work
  8. work-liveresearchbench-61e7da557bLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — examined-source-work
  9. work-reportbench-dca823c910ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — examined-source-work
  10. work-researcherbench-82729ab60fResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — examined-source-work
  11. work-url-health-f64342d493Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — examined-source-work
14 · Examined editions
  1. edition-citation-verifier-arxiv-v1-acfd4abab7Quote-minimal projection of edition:citation-verifier-arxiv-v1 — examined-edition
  2. edition-cnv-arxiv-v1-f24001c50aQuote-minimal projection of edition:cnv-arxiv-v1 — examined-edition
  3. edition-deeptrace-arxiv-v1-70a6e720c1Quote-minimal projection of edition:deeptrace-arxiv-v1 — examined-edition
  4. edition-deeptrace-iclr-2026-b2b153eaa6Quote-minimal projection of edition:deeptrace-iclr-2026 — examined-edition
  5. edition-drbench-arxiv-v1-7afa5ade36Quote-minimal projection of edition:drbench-arxiv-v1 — examined-edition
  6. edition-drbench-iclr-2026-d00a3dcdb6Quote-minimal projection of edition:drbench-iclr-2026 — examined-edition
  7. edition-drbench-repo-main-469cce5-4f8964f3e3Quote-minimal projection of edition:drbench-repo-main-469cce5 — examined-edition
  8. edition-keplinger-supplement-v1-f6ddb4c04aQuote-minimal projection of edition:keplinger-supplement-v1 — examined-edition
  9. edition-keplinger-supplement-v2-01e44c09d8Quote-minimal projection of edition:keplinger-supplement-v2 — examined-edition
  10. edition-keplinger-vor-2025-10ba10f860Quote-minimal projection of edition:keplinger-vor-2025 — examined-edition
  11. edition-liveresearchbench-arxiv-v5-f5644763e5Quote-minimal projection of edition:liveresearchbench-arxiv-v5 — examined-edition
  12. edition-reportbench-arxiv-v1-25888dc44bQuote-minimal projection of edition:reportbench-arxiv-v1 — examined-edition
  13. edition-researcherbench-arxiv-v1-39a8b256d7Quote-minimal projection of edition:researcherbench-arxiv-v1 — examined-edition
  14. edition-url-health-arxiv-v1-85fb934b53Quote-minimal projection of edition:url-health-arxiv-v1 — examined-edition
72 · Accepted exact span roots
  1. span-015bb68c904bf225e84c82b1acffb3cc38a4a087d7b0f79d-4c454092b1Evaluation methodology, lines 144-149 — matched-exact-span
  2. span-084f11e57cf779850ea87aaf30ba1277438ff9efcb533fd2-fe72b27cfbAppendix A.1 'Limitations' — matched-exact-span
  3. span-092745c7db6407d4d52927642741cde229fad05a5347bb18-7b04227f8fMain text, paragraph beginning 'For their high performances in generating reference lists' — matched-exact-span
  4. span-0abbab3d6269d58eb76d8e90eff1e4c27512cf46d9b8e5ff-54e5759a8brepository README, Overview — matched-exact-span
  5. span-0d4f357f2a5565bdaea01a1e537650958d71216868d1a2b0-9b0af35d76Table 1, lines 177-182 — matched-exact-span
  6. span-13cb0ddb087b82988b7445a9cfb02562f9c044ef4f244e8b-87c3b42c68Section 3.2, lines 144-146 — matched-exact-span
  7. span-196c54d6212fbc0c054dbf8e8467f2388c1f788bb93f4b1a-d460e927f6Table 1, Perplexity Deep Research citation-accuracy cell — matched-exact-span
  8. span-1db12b68cfeb4c5f5a96fa9cd2e2bd46e8dfb4f096f73154-cea47f0c65Limitations, lines 259-261 — matched-exact-span
  9. span-279db725380c133044944dccb379b75a396242219aa30437-569fc9e2fdpage 16, FACT judge validation — matched-exact-span
  10. span-2f311bfb451c6208101cb7583af37c338914fcdb10f5c75b-90e7764ce2Table 1, OpenAI Deep Research row; FACT columns — matched-exact-span
  11. span-3015a109e7898cb720e426e9a57c526277b648c75f1fb4b1-745d84ff44Section 4, line 186 — matched-exact-span
  12. span-35927851edd2e7e5e0f60d98498940f88304ba99bf1f85a0-2fe586436bpage 14, Table 3 — matched-exact-span
  13. span-39a9931f05ea356fe904ed15768fa08c66440b1fcb20f98e-191c24bbb2Table 1, ChatGPT all-correct subtotal — matched-exact-span
  14. span-417182b4c15ac550469e435a02aab1e91b395f72f186312d-e9e2c4a70cSection 4.3, line 212 — matched-exact-span
  15. span-46c392bc942cd88d525d74298e6bc6ebd9aeaec533ddfb7b-774f5b996emain text, paragraph beginning 'For their high performances'; Figure 1 — matched-exact-span
  16. span-55de6a6ca9414a252592e9f77544a9259533dd713ce1f12f-0826216ffcpp. 1-2, lines 64-95 — matched-exact-span
  17. span-59d0da64f738e1d8745f2bff3663eb89e273f93fe501afe3-ba69fc58ebSection 2 definition discussion and Section 3.3 — matched-exact-span
  18. span-5cd9ba4ca72da10cf51c0beddc6ea2765c74721bdd526a67-ed8f789629Limitations, lines 219-224 — matched-exact-span
  19. span-6885cf2524462022475b5829da081b2d35fcfe7463ca6ad0-3b1a122fafSection 1 contributions — matched-exact-span
  20. span-6c1d327e58ef001533ea1c11010e974436bbeafbde138efc-f6efd8d376Section 3.3, line 155 — matched-exact-span
  21. span-6e961498ed0f81a950f775956587f741e0eb9f11640fd3cb-afa5b59854Section 4.3, Tables 2–3 — matched-exact-span
  22. span-6f79c3c824ccf889d73e234335661e5fca68d90f8c78ee97-dbf7848a72dataset page, Description — matched-exact-span
  23. span-732c03e69c6e142a492275cee5d1e2d474c75422c9b84aaa-f7675998ccSection 3.3.1-3.3.3 — matched-exact-span
  24. span-74f6acafbb4c54d556a572460bf1b45bacd1e293f4aa16cc-5ee15c82c4Abstract, lines 47-48 — matched-exact-span
  25. span-76e9bd59b31c7ebd2ed7a2186e59753a79d3c730e0d7b4fb-9c98d99242Appendix D, Table 5 — matched-exact-span
  26. span-794bdf7e0595799afbb0d55f3dc522077aed9615c59f1e35-43f2ba944dAbstract, line 16 — matched-exact-span
  27. span-7a6046a756d6ec0315d330cf0817d062ea48498bdc9c90b3-c3a1e80029p. 5, lines 315-329 — matched-exact-span
  28. span-7d7331942f9d0520e18e96206e7718e5187788be4d48cef9-dfdbc1d019page 4, Section 3.2, Support Judgment — matched-exact-span
  29. span-886b50d94e3facabc487342cd2af5157cf956622ecfc3686-e7bdad0935Section 4.2.1, Citation Support Verification and Score Computation — matched-exact-span
  30. span-8c8d9a6dc003bf0298542b85df3f6f08db631e68e89f606f-92731379f2Table 1, ChatGPT subtotal — matched-exact-span
  31. span-8cce3d9ee1e1674ba1b93321dd9b99a55d834ca980451f8a-d2fceafe11Abstract, line 46 — matched-exact-span
  32. span-912c931063109f84ea88aea34891dd5e0f9147d2176af258-4d566ca076Abstract and Table 1 — matched-exact-span
  33. span-92eb7cf19745e24190dda84c77d590941134d09f06805c0f-88afe43377README FACT, lines 240-248 — matched-exact-span
  34. span-94ad583514aa7e04b8d5f7b8f9cdda6fdffe8ca6712f033d-63c1ba24ebDataset description — matched-exact-span
  35. span-9c7a4c186cc85f185aa293a7c5a46c08f9dbbc1a9e63c449-feff07a6e2Section 2.2 'Cited statements' — matched-exact-span
  36. span-9cb1700c08368feb234207045354cbf559d6ed80f2b0bae8-f898722eb5pp.4-7, Sections 3.1.1 and 3.2 — matched-exact-span
  37. span-9dd6eb46529e77a62748093528e52685fbd92d46a756aef7-96d75d9e2bSection 3.2, lines 163-171 — matched-exact-span
  38. span-9fe42cb48702d8d96ebbbc0309692214ece6ec495c447c06-80da541506Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment' — matched-exact-span
  39. span-a08468cb62f2feb21d9cbb95175ce00f54592bd83573cc91-03f65cfa45Official ICLR abstract — matched-exact-span
  40. span-a15ab02e9fcb23df02fb03d28967d4220ddf65ad380d2230-ddf85fa359page 4, source scraping — matched-exact-span
  41. span-a18e778fac3d49e81167f05e09fbc361e91af8c0c01b0fec-ca1b18e206page 3, Section 3.3 — matched-exact-span
  42. span-aa1534ce2e177153ab85e512deb8d064d14ec1eae16d8032-2299096a09p.14, Table 3 — matched-exact-span
  43. span-b0a39c963cdfdc5f9852cfc9f7e33e05d31232bd1d899885-adf9545222p. 4, lines 150-159 — matched-exact-span
  44. span-b38a2af5982b4c85bc205d2f533a23ed3a8f40a49b651a5f-27395a6291Main text, paragraph beginning 'To address the gap' (search-indexed primary full text) — matched-exact-span
  45. span-b68153666e4e12aee8dc887fef0a61fa5f44b5f680a0b7eb-769c2cc17dSection 5.1 Results — matched-exact-span
  46. span-b6a6ba48a2c4e352c8565b671a15a90400739f2aeefda8a6-e8232daf27Main text, following paragraph (search-indexed primary full text) — matched-exact-span
  47. span-b83996e3b5fdf878e04d6d41d0e7a1eebee9fdd9bbccf2ae-820a3eab86Appendix C, line 365 — matched-exact-span
  48. span-b903dc2284e1b4d80bb6ef948d79bbdbc7d7447e871281b8-6ef6234316body, paragraph immediately after Table 1 — matched-exact-span
  49. span-ba6b8a8b7d9121ebb05d594176ceed3d1e9dbc637876783a-d85147b6b0Settings and metrics, lines 132-139 — matched-exact-span
  50. span-bb16d5f06abe8641d9ea6694e549aa15e240dd88946aa47a-5946b6e6b7Abstract and Table 1 — matched-exact-span
  51. span-bb62928b5e0cb4e373b2f6bfeb579d8564d05da5588f8b15-771310f4efSection 3.3 'URL extraction and classification' — matched-exact-span
  52. span-c27aaca02bda06ed764be53351158fc862af9d4a556d3e82-104e7cda52page 5, Section 4.1 — matched-exact-span
  53. span-c3dfc98566c27561eb7267972e00e3e1ba6a399a3ff5d002-37f40d3260page 5, Section 3.3.3 — matched-exact-span
  54. span-c7dbd1d9e09980703ecbee8161594e4b426c1af63117b286-153d26f046Section 5 'Limitations' — matched-exact-span
  55. span-cd21c450c0d33df20a2b1a539d5725347db1b5bab7d1e986-2e3e0c3ac7Abstract, lines 45-49 — matched-exact-span
  56. span-d638c8b21f0e4a0f840880b6ffe0217ce1623758cc09fdcc-8585f6d39bSection 4.3, Tables 2-3 and following paragraph — matched-exact-span
  57. span-d880991ad784390029d49f84949d083c4789d67d45893d87-c6ef8a5c13p.33, Appendix E — matched-exact-span
  58. span-ddf4320a6d302c13d28b6225eb7e79ac15583349e1b449d4-f89bb46179Section 3.1 'Setttings' — matched-exact-span
  59. span-deada8bacb987a47f2a42e52fbfeb5c1beb3d250c0be79e2-0e91f5caccSection 4.4, line 217 — matched-exact-span
  60. span-e223640cb0ed12c87ca1d2406f3276f30a3b8d2017dd4dc1-3cfecd76bfTable 1, OpenAI Deep Research row and cited-statements Match Rate column — matched-exact-span
  61. span-e2c9f6220843a04665f5d1cd142ac15f04211ace5c579fa6-c5e0476df0Paper section 3.1.1, factual-support validation — matched-exact-span
  62. span-e4a9f30effc031388ed5c9993dc6f071903bec5e3f1998db-5d95269b76pp.17-18, Appendix E, equations 4-6 — matched-exact-span
  63. span-e54ee1457854a9a5e46f7c7c2e66ac6a5a69b963f8ee3450-8f875cfe94Abstract, line 16 — matched-exact-span
  64. span-e5bcf0241a1aeae7d9a92755b1bcf0adeeea93a8922f22f8-77a5ff4cc0README News, 2026-05-11 evaluator migration notice — matched-exact-span
  65. span-e6f447129d9a662fb202d4ac22e16822bacfdbcbcdb7afe3-f7abe2194cp.23, Appendix C, Citation Accuracy — matched-exact-span
  66. span-ea9cb2e1b0078c0857b789d61d0a0d26801cf66891864705-2f19633d83Section 5.1 and Appendix E, Table 4 — matched-exact-span
  67. span-eaf37bef7b8dc35f3e00b85d3e58ad872dc0ea604a062c8c-fa09300f85p.4, Section 3.2, lines 148-155 — matched-exact-span
  68. span-ef637b38f36195bea09205d532e8744daf3e257f93d97144-26007673a3README News, lines 176-188 — matched-exact-span
  69. span-f05b3189dc3fafd45b5bde643a91fd0616832928215a37a2-47f72e77e4Paragraph after Table 1 — matched-exact-span
  70. span-f35c9565f2fd047f6e7272dc608bbf6a63813a717bad5354-e6a07f05b6Appendix C, line 365 — matched-exact-span
  71. span-f96ad5daa81153a3689f52f5777fedfd70e8a87d44e7a1c9-3da72b8748Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3) — matched-exact-span
  72. span-f98f1274859ba98cc538e3be6d11d345a5382cfcb9516aa1-d60b765c48p.17, Appendix C, lines 704-708 — matched-exact-span
7 · Candidate warrant roots
  1. relation-cnv-link-relevance-support-gap-165443d270Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.
  2. relation-cnv-search-depth-ablation-9c01321092Within the Cited but Not Verified harness, the 2-to-150-call ablation reduced reported Fact Check scores for two setups without establishing a general causal law about search depth.
  3. relation-deepresearch-bench-fact-a3f61d7a3aDeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.
  4. relation-deeptrace-support-variation-1a84ef7c6fDeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.
  5. relation-keplinger-metadata-versus-support-c23e105c8bIn one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.
  6. relation-reportbench-match-rate-6964e41a69ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.
  7. relation-url-health-resolution-0b26fb45c7The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.
0 · Independently confirmed warrant roots
    4 · Pending warrant groups
    1. relation-citation-verifier-calibration-5b40082faaThe captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    2. relation-liveresearchbench-e1-e2-e3-5fb5d671f7The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    3. relation-researcherbench-faithfulness-groundedness-c84dcaee54The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    4. relation-url-health-correction-loop-eb8a6384c6The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    34 · Unresolved citation occurrences
    1. V2-SOL-01:s1_keplingerAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    2. V2-SOL-01:s1a_keplinger_dataSupplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    3. V2-SOL-01:s2_drbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    4. V2-SOL-01:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    5. V2-SOL-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    6. V2-SOL-01:s5_url_healthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    7. V2-SOL-02:S1Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    8. V2-SOL-02:S3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    9. V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    10. V2-SOL-02:S5ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    11. V2-SOL-02:S7Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    12. V2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    13. V2-SOL-03:s3DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    14. V2-SOL-03:s4Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    15. V2-SOL-04:S6Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    16. V2-TERRA-01:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    17. V2-TERRA-01:s2_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    18. V2-TERRA-01:s3_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    19. V2-TERRA-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    20. V2-TERRA-01:s5_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    21. V2-TERRA-02:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    22. V2-TERRA-02:s2_liveresearchbenchLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — unresolved
    23. V2-TERRA-02:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    24. V2-TERRA-02:s4_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    25. V2-TERRA-03:S1Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    26. V2-TERRA-03:S2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    27. V2-TERRA-03:S3ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    28. V2-TERRA-03:S4Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    29. V2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    30. V2-TERRA-04:S1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    31. V2-TERRA-04:S2_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    32. V2-TERRA-04:S3_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    33. V2-TERRA-04:S4_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    34. V2-TERRA-04:S5_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    20 · Unsupported or force-raised claims
    1. V2-SOL-01:r3_deepresearch_bench_factV2-SOL-01:r3_deepresearch_bench_fact — no-credit
    2. V2-SOL-01:r4_deeptrace_supportV2-SOL-01:r4_deeptrace_support — no-credit
    3. V2-SOL-02:R3_deepresearch_bench_factV2-SOL-02:R3_deepresearch_bench_fact — no-credit
    4. V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — no-credit
    5. V2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — no-credit
    6. V2-SOL-03:r4_deeptrace_supportV2-SOL-03:r4_deeptrace_support — no-credit
    7. V2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — no-credit
    8. V2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — no-credit
    9. V2-SOL-04:R1V2-SOL-04:R1 — no-credit
    10. V2-SOL-04:R2V2-SOL-04:R2 — no-credit
    11. V2-SOL-04:R3V2-SOL-04:R3 — no-credit
    12. V2-SOL-04:R4V2-SOL-04:R4 — no-credit
    13. V2-SOL-04:R5V2-SOL-04:R5 — no-credit
    14. V2-SOL-04:R8V2-SOL-04:R8 — no-credit
    15. V2-TERRA-01:r1_deepresearchbench_factV2-TERRA-01:r1_deepresearchbench_fact — no-credit
    16. V2-TERRA-01:r4_cited_not_verified_source_attributionV2-TERRA-01:r4_cited_not_verified_source_attribution — no-credit
    17. V2-TERRA-02:answerV2-TERRA-02:answer — no-credit
    18. V2-TERRA-02:r1_deepresearchbench_factV2-TERRA-02:r1_deepresearchbench_fact — no-credit
    19. V2-TERRA-02:r4_deeptraceV2-TERRA-02:r4_deeptrace — no-credit
    20. V2-TERRA-03:R3V2-TERRA-03:R3 — no-credit
    9 · Independently rejected claims
    1. V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — independently-rejected
    2. V2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — independently-rejected
    3. V2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — independently-rejected
    4. V2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — independently-rejected
    5. V2-SOL-04:R1V2-SOL-04:R1 — independently-rejected
    6. V2-SOL-04:R2V2-SOL-04:R2 — independently-rejected
    7. V2-SOL-04:R3V2-SOL-04:R3 — independently-rejected
    8. V2-SOL-04:R4V2-SOL-04:R4 — independently-rejected
    9. V2-SOL-04:R8V2-SOL-04:R8 — independently-rejected
    3 · Inaccessible carriers
    1. V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    2. V2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    3. V2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved

    Five-line lexicon

    The units are different on purpose

    Citation occurrence
    One report-level use of a citation, before repeated URLs are collapsed.
    URL root
    One distinct cited URL string; resolution does not prove claim support.
    Source work
    One logical paper, dataset, repository, or audit instrument across editions.
    Candidate warrant root
    One source-method-data proposition that survived bounded semantic review, while residual independence remains unresolved.
    No credit
    The item remains visible but does not support the stronger claim because its carrier, span, semantics, or lineage did not close.

    Build receipt

    Reproduce this projection

    Reproducible projection
    Dossier
    em:dossier:sha256:cbd7a14096a956f642f5c76046d3b49ed648fbe6bf24144c992404a01415af82
    View policy
    em:application-policy:skeptical-v0.1
    Catalog
    em:catalog:sha256:092898e1fe3d355761ab4cec653576926a8f5d31621ec7ce23dd60e9d19563ef
    Frontier
    em:frontier:sha256:7e4a173112ef26422acf3ed9434c8b6849c4e011797e20fed6c0a9ca58a1e4c3
    Accepted commit
    af081caa99fc08d3fabb914ff68f2e672a83bd5b
    Epistemic policy
    commons-balanced-v0.1
    Disclosure policy
    public-noninterference-v0.1
    Compiler
    epistemedia/0.2.0
    Content digest
    3e25148910d8bf697cec3ed8bc17eefdbb6638d68d4b3f4b75700c8b95c1cbd7