How We Know · Case 002 · encyclopedia

When eight research agents agree, how many evidence roots are there?

What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?

Eight context-isolated public-by-design reports captured under one frozen prompt, plus the public source editions and exact spans they cited. This historical pilot does not estimate current or universal agent behavior.

encyclopedia finding

The bounded record contains empirical methods for URL resolution and claim-to-source support, but the eight reports reuse overlapping capture, source, method, and derivation lineages. Report agreement alone adds no independent warrant beyond the inspected source record.

What to do with that: Use a polished cited report as a map into sources, not as a vote count. Check URL resolution and sentence-level support separately.

This bounded 2026 packet does not estimate current or universal agent reliability.

Lineage accounting

Agreement is not a vote count

Shared capture: Unknown provider and retrieval dependencies remain; all reports share the exact prompt and one bounded capture program, so run multiplicity gets zero automatic independence credit.

Warrant boundary: Unknown residual independence remains across task data, judge methods, retrieval, source, edition, span, derivation, and upstream-citation lineages.

Sentence x-ray

What the selected record actually supports

Open a sentence to inspect its exact work, edition, span, retrieval, digest, and license chain.

01 This report is a captured observation, not an independent evidence root.

Typed relation: dependence

EM-0026 deterministic agent-citation evidence ledger

Captured report V2-SOL-01

{ "answer_bytes": 40476, "answer_path": "research/how-we-know/agent-citation-lineage/answers-v2/V2-SOL-01.json", "answer_sha256": "17a7cf97b717d7b520026b9eb52a7da59ddb9b5f698c2aeadfab3f38c91008a0", "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_model_profile": "gpt-5.6-sol", "retrieval_infrastructure": "unknown", "run_id": "V2-SOL-01", "status": "completed", "trace_path": "research/how-we-know/agent-citation-lineage/traces-v2/V2-SOL-01.json" }
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:ead238d33280cfd1dd252936d47fc1351950f8efb4105085d161ac7c1e51fa97
Span digest
sha256:ab05cfa54280c980dabd52b0d336f1dc65a5761925a54dad2d5c7b1cbfb094cb
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments

EM-0026 deterministic agent-citation evidence ledger

Shared capture lineage

{ "automatic_independence_credit": 0, "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_profiles": [ "gpt-5.6-sol", "gpt-5.6-terra" ], "retrieval_infrastructure": "unknown" }
Edition
em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
Edition digest
sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
Span
em:dossier-span:sha256:e120064fee0bdabe15299cb865ff2f7df66473291a13b7a765a4a09d16223f5e
Span digest
sha256:92b558b23559a96a551de8acb2627b148688d21ede2ca3fb9966196c4675dab6
Retrieval
No external retrieval record; repository audit span
License
Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
02 In one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.

Typed relation: support

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Main text, paragraph beginning 'For their high performances in generating reference lists'

citation-bearing sentences exhibited high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:9f91508aa7be57f8e16c902e5399763ce349e7318b10d6085b65f0f53d5bdb56
Span digest
sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Table 1, ChatGPT all-correct subtotal

16 (69.6%)
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:72a3edb3976aa397d38e9a7ccbc18b5e4af6f00c91540cd47aadcf9eddce2439
Span digest
sha256:b6c3ff6f2ff60f37ae1c6b773ba6aa5a6c01015b61e3cc6f2a06bfcef69b53ad
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

main text, paragraph beginning 'For their high performances'; Figure 1

citation-bearing sentences exhibited high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:910f7758be3883fc0964cd1e5412b9c75a586e6c4c7c08e64ebda5b9f541a132
Span digest
sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Table 1, ChatGPT subtotal

23 (100.0%)
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:d9dab2e8e8da55f5a5ddbe3e62c658a7f643cd05c36c117aa8a0dfc76a7b6d09
Span digest
sha256:5428bd1a04d5d2e332015107e915fe3a38f1d5e173e8c7246b743d98ecfd8618
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype

Dataset description

sentence-level evaluation
Edition
em:dossier-edition:sha256:8c80a50cd40cbb31bdaedb6730e8e173e3edc0f3440538dbcc4878eb0997ffd9
Edition digest
sha256:0a485e13dcc56e5f883213fde7e24bdcc4ffb810bda8ac95186f83f6325fafd1
Span
em:dossier-span:sha256:38bf64d13c4edef9c0aba131c84d42e0e08637a661d2552cbf2cb49536813488
Span digest
sha256:e3ec5c81f615c657f7c6b124c5b2150d86f18281384ca0bd24ebb43a3317796f
Retrieval
retrieved https://data.mendeley.com/datasets/3s73z9zf3c/1
License
CC BY 4.0 · metadata and quote-minimal landing-page spans

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Main text, paragraph beginning 'To address the gap' (search-indexed primary full text)

We imposed a 1500-word limit, APA-style references and conducted three independent runs.
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:54d8710861e7e1ee03c50697f1316c12138246e303a08c843fdadecfa99b8c5e
Span digest
sha256:1c94be957cfe750f445dc9143722905b5784b02f754a9b4071a3381e2bd493e9
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Main text, following paragraph (search-indexed primary full text)

ChatGPT Deep Research-generated references were mostly identifiable by their metadata (95.7%): 69.6% were entirely correct
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:b9e4b58f255c32e2a9e5e6796e8d35f7f28d7481060d6b35f1e02a3c74d2b3d6
Span digest
sha256:28141fd1b0222658eb93a338783358926de770223eea24f8c5cbcf632d67f955
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

body, paragraph immediately after Table 1

high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:f687ae7c3889980d4d3eea1b63b7bea25ae674ebc04caf8c59e7ae4f4fb77464
Span digest
sha256:2e43a0652065a6a1fb69de224082ac79d903cc71cac0664a1cb376cbfa09d0e8
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype

Paragraph after Table 1

51.3 ± 6.5%
Edition
em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
Edition digest
sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
Span
em:dossier-span:sha256:2f331f1e2636907b0cab1000b0947e440fa5f24eec0edaf9498470522e1abe1a
Span digest
sha256:1f863c1f35c37af1fd2c3e4927ea0b289d94a9532fa5855e209ce98ea344d498
Retrieval
retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
License
CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
03 DeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.

Typed relation: support

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Evaluation methodology, lines 144-149

Each unique Statement-URL pair undergoes a support evaluation.
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:ed08815f4aabd455caf37f61a1e37a2dc35725098f2e9be5c6da6c9d5854d612
Span digest
sha256:1bde849b0f5df91e847e907dd763f27be01ffdc7eb0c7e61898c8a9731472c35
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Table 1, lines 177-182

Perplexity Deep Research | ... | 90.24 | 31.26 ... Gemini-2.5-Pro Deep Research | ... | 81.44 | 111.21 ... OpenAI Deep Research | ... | 77.96 | 40.79
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:fd8835d2aaa9358bd7dfbcf07848ed19001fa82f11e45ceff3374cdd4e882be3
Span digest
sha256:53d38f9cf0d0aedf28a77d7de32935372202fff78a0c02d98f8c269ef223906a
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Section 3.2, lines 144-146

‘support’ or ‘not support’
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:4648142e8019a2d54f201844685fbb87a921d93fc24b0f5c08f8124a82ffc374
Span digest
sha256:a49c200ea10833496ac73719992be09d5fd3fd021fdd51ccc587ac2592225f56
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Table 1, Perplexity Deep Research citation-accuracy cell

90.24
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:b118f44725ec505886f92a94123fe289be763523b8afb669757c72ccd0d8c352
Span digest
sha256:64c911242980826f92ffae6d0571646ecc3aceca73c213791c6d3a8d889b4cc7
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

page 16, FACT judge validation

aligned with human 'support' determinations in 96% of cases
Edition
em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
Edition digest
sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
Span
em:dossier-span:sha256:a59ed2a9ea0e6972bc3f71d63feb1c361b76eb8d8fa5bb679c7beab61a32b713
Span digest
sha256:8be6257e74c3696813fb0079c27e93e8bec16c71d214b6da43811fbd9b406ba4
Retrieval
inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
License
CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Table 1, OpenAI Deep Research row; FACT columns

OpenAI Deep Research | 46.98 | 46.87 | 45.25 | 49.27 | 47.14 | 77.96 | 40.79
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:762ebe0afa7a2ddf8eb19d879c91f63d1de7d25192f10d248b65fefb0e113038
Span digest
sha256:dedd156999fa7d57ae9c22a166c56c263547eee3c37e6f41fe7644d980110986
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

page 4, Section 3.2, Support Judgment

This yields a binary judgment: 'support' or 'not support'.
Edition
em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
Edition digest
sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
Span
em:dossier-span:sha256:d915faef70b54c6751577ba10f4057744b2b1b0b4d10613cdd8d74720a5da23b
Span digest
sha256:78f021c07980218ead835a986c84780444479e627dcfa9f2ea126b3604b89ad7
Retrieval
inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
License
CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

Ayanami0730/deep_research_bench

README FACT, lines 240-248

Support Verification: Uses web scraping and LLM judgment to verify whether cited sources actually support the claims
Edition
em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
Edition digest
sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
Span
em:dossier-span:sha256:a28651fed7ffef6fdbbe7f322faa2482aa1e0f6b1e2a0a909eddef59a72740f2
Span digest
sha256:549018c6df261f27bd10133ad4e4f565e004f07710095259a7889cc79d149794
Retrieval
retrieved https://github.com/Ayanami0730/deep_research_bench
License
Apache-2.0 · metadata and quote-minimal README spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment'

"Each unique Statement-URL pair undergoes a support evaluation."
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:aba61c5b1685e20da29c0d6a42dd49b66f0b6c2f3ca08483716ae5c518562d4f
Span digest
sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

p. 4, lines 150-159

"Each unique Statement-URL pair undergoes a support evaluation... This results in a binary judgment ('support' or 'not support') for each pair, determining whether the citation accurately grounds the claim."
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:0b565d8281beb5c9b95d300108446049bc4c93de4dc3e80a501f41e85d0a98a5
Span digest
sha256:4edb75ec3f690eb2d15089923ba5e9c4e6174f55094d0919041bf10625901ceb
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Appendix C, line 365

96% of cases
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:d753bc9c860da325190141f3f78a0d031f2d6c7cf5e1a51f19f355f44cfb47db
Span digest
sha256:ae845f85955fa47d52d51f32803e6a26afa8b725c523e53d9ef30b2abca72d9b
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

page 5, Section 4.1

complete set of 100 tasks
Edition
em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
Edition digest
sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
Span
em:dossier-span:sha256:4aa9441f93251fa969f1fe8cc8a4fb6cf728ec3f8879efae3ffff55fabae759d
Span digest
sha256:11034cf2238003f213bf87210c9017f57cc743910fa70754f3e61f875abe75e6
Retrieval
inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
License
CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

pp.17-18, Appendix E, equations 4-6

"the proportion of 'support' statement-URL pairs for each individual task"
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:c6710bb4b411353c7c600e449cc556e39a3530dc63c002a7b3973ef6bb1793bf
Span digest
sha256:82b9361b6e696a92b41fdfae500f56b05d48575c0d429f15b3a8d6786c83b8d7
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

Ayanami0730/deep_research_bench

README News, 2026-05-11 evaluator migration notice

Legacy code: the previous Gemini-2.5-Pro / Gemini-2.5-Flash evaluation code is preserved on the Gemini-2.5 branch.
Edition
em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
Edition digest
sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
Span
em:dossier-span:sha256:abc4ab811e5c574575a0024557dda57a45d8d439ad4ff4ce08bef3753f7a86ef
Span digest
sha256:02cd1c04d6312fcda76549885095d9c9d971cc14709ae8f749cf47fc2c11bcee
Retrieval
retrieved https://github.com/Ayanami0730/deep_research_bench
License
Apache-2.0 · metadata and quote-minimal README spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

p.4, Section 3.2, lines 148-155

"Each unique Statement-URL pair undergoes a support evaluation."
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:27d667cb8cff21b646ec759c6c6fe79c5744a28079337d7a7fa9ac98abf8da9d
Span digest
sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans

Ayanami0730/deep_research_bench

README News, lines 176-188

Official Evaluator Switched to GPT-5.5 ... Legacy code: the previous Gemini-2.5-Pro / Gemini-2.5-Flash evaluation code is preserved
Edition
em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
Edition digest
sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
Span
em:dossier-span:sha256:fd761963efd1f957ebe8dae1eb76239c2d17ce1553109a4e0ac420525c73cfcb
Span digest
sha256:9c713be4fc872b46e66d812e033708b51b4e385c81aa45f7b1798459a677343d
Retrieval
retrieved https://github.com/Ayanami0730/deep_research_bench
License
Apache-2.0 · metadata and quote-minimal README spans

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

p.17, Appendix C, lines 704-708

"aligned with human 'support' determinations in 96% of cases"
Edition
em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
Edition digest
sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
Span
em:dossier-span:sha256:17193f2aab7106eece041bcabda963d9bd172d798792efc058b6a93c9bd4cb0c
Span digest
sha256:43c892b582d42a335c982009f97347b5bf2ae2799cc3e478c53504f4d01e4388
Retrieval
retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
License
CC BY 4.0; unknown · quote-minimal attributed spans
04 DeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.

Typed relation: support

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

page 14, Table 3

Factual support (statement-source) 0.62 binary
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:6cf2883b02dd404844a662bae3367ab60d91f6847160a3df9c6ec6b287dc0858
Span digest
sha256:bb4f28cd0f4527c633bcfacf51b18ced7387c4a2f266c59d41de7ee899b24b85
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Abstract, line 46

citation accuracy ranging from 40–80% across systems
Edition
em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b
Edition digest
sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d
Span
em:dossier-span:sha256:e748230bbc2f6b31433eaba4536917834c454cb73a7f18e8d9b975e2cac447b5
Span digest
sha256:a0eeca5c86f9c7aeda9167f626a8cb0ed56dc87ea1ac5ca8c4a7a0add76e7625
Retrieval
retrieved https://arxiv.org/html/2509.04499 · retrieved https://arxiv.org/html/2509.04499v1
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

pp.4-7, Sections 3.1.1 and 3.2

"The dataset comprises 303 questions"
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:b9cbe6b1d35e74f3c8133d2be41a47d9361a26958e9f644dafd97e9e5acb516a
Span digest
sha256:2ce91759797b86455a478a9ced58ec59d032e1a32a1c964347f96be4ab9e2054
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Section 3.2, lines 163-171

303 queries x 9 models
Edition
em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b
Edition digest
sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d
Span
em:dossier-span:sha256:8c1ac40e63a414c1a41cd0842632677269c8800967f22ef8247b6fb36f6242ea
Span digest
sha256:5cc83ab699cb1342475b407e34d33fc106d64b5d9270e5bb9445d80586ebabc2
Retrieval
retrieved https://arxiv.org/html/2509.04499 · retrieved https://arxiv.org/html/2509.04499v1
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Official ICLR abstract

with citation accuracy ranging from 40–80% across systems
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:06ea026aae360f700369f4b1f658c4ca85e8acce431923cd7978ff2358b1ec05
Span digest
sha256:af76864a50c641f8d693216d3b7ebe2a414eebf75467305c1e1a6946b6746590
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

page 4, source scraping

For roughly 15% of the URLs, the Reader tool returns an error
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:12ab62355d178256bbaed6cdf5cbf6ae150410d9364f98cced347b54161fa5e8
Span digest
sha256:a38c9ac275570f2b78dd1e1bf7663579e35f9037131e12d721dfb4c6f005683d
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

p.14, Table 3

"Factual support (statement–source) 0.62"
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:2aa386b3653a99db3b670f783272056371480ed287dc1f62a51d3d587d680996
Span digest
sha256:75a9212dc448c15ffc1e4a36bdf190728977470661bd9f45be534010704faa25
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Paper section 3.1.1, factual-support validation

Pearson correlation of 0.62 between the LLM judge and manual labels
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:a575506a3cca26acda376aa922fdb0e269767f71b01f502a7c50f3d4492e6f03
Span digest
sha256:d2f6a4561a938ee913dadde2fc4a80c32c492fbcb3b47f236c9639375843ae8d
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

printed page 10; PDF page index 10/22; Section 4 Results; Deep Research Agents paragraph

Gemini(DR) demonstrates weak citation performance: only 40.3% citation accuracy
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:533550e262b8691ce9ed33aa0e89a7bf13d89c04c6002e1855ed1e767f556e3a
Span digest
sha256:cf32a686cec566152749e3a4c643fa29aa115099f9eba77c0ca113c48264f769
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

printed page 9; PDF page index 9/22; Table 1; Gemini (DR) column; %Citation Accuracy row

{ "cell": "50.3", "column": "Gemini (DR)", "row": "%Citation Accuracy" }
Edition
em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
Edition digest
sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
Span
em:dossier-span:sha256:cf4e063e6f297fe901b342752b67850d84cc3e8e21ace2ef8d5ea4f3fd93091c
Span digest
sha256:e020a09063eee7aa1eb70663cf3e8bd7a8162952f5ae65d959d57ef74ab855f6
Retrieval
retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
License
arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
05 Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.

Typed relation: support

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Section 3.3.1-3.3.3

"Fact Check verifies whether specific factual claims are accurately supported by the source content."
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:d4e2f35392b064365ad867f33ead8636bb8bbf3af03c0141585e5ae56c868090
Span digest
sha256:a4386d56325583893d806fdd15ccd475d5096f0b445b1b9f52aeba253c910987
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Abstract, lines 47-48

yet achieve only 39–77% factual accuracy
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:ea33bc17365ed01ed68ab06f33c4fd2020dc8e6947659e4e9629b44af4aef635
Span digest
sha256:be05510d58cc0231bf968f29c70b3d669a82d5cd0c74a778a5f8c4e035bf45b4
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Abstract and Table 1

link validity above 94% and relevance above 80%, yet achieve only 39–77% factual accuracy
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:49ee66722d6d523dd485f05c796181af55736c3ceab2a235c5c3b40cca95e439
Span digest
sha256:44d4c6941f2920e907de1e4cd22aeb9a725f24cc19432723754e316bd6f5cc49
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Abstract and Table 1

maintain link validity above 94% and relevance above 80%, yet achieve only 39–77% factual accuracy
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:8cf2ff4c32b2dac6d6dfc7684dd2bf7dca9d40fe7b7b36d50dfdcddad067bff3
Span digest
sha256:e35ca66d12349d7290a6dbe7449bf6da173eb0671765644b581713528b316c74
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Abstract, lines 45-49

Citations are evaluated along three dimensions. (1) Link Works verifies URL accessibility, (2) Relevant Content measures topical alignment, and (3) Fact Check validates factual accuracy against source content.
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:237508b682de825711f61762deb5ce7661d5d197e9fde3aed5f5e80f05fdcdb5
Span digest
sha256:d34096a9d9b1fbb43ef50ba503c3894607c008b563e1e1a5d1643cbc23c5d447
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

Section 4.4, line 217

only 1 failed link out of 2,159 evaluations
Edition
em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
Edition digest
sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
Span
em:dossier-span:sha256:2bd67a5caa606894f4301ca1185c88a4dde02d3c5c4a2cf311463a0f39332443
Span digest
sha256:da34a8e2a1d4e190d3b2095d79886b3cbf7edb2c63874189aa381302d2c2924e
Retrieval
retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
License
CC BY 4.0 · quote-minimal attributed spans
06 The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.

Typed relation: support

Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents

pp. 1-2, lines 64-95

"This study focuses on URL-based citation hallucinations"; "fabricated snippets ... and invented bibliographic entries ... require separate systematic study."
Edition
em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
Edition digest
sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
Span
em:dossier-span:sha256:179697ec62048a6465fa50dda34ff54530e700430013a878201b9dadde405cc0
Span digest
sha256:ee1de7374c67e208aa5d1331c20a08c8c83418309f1bd768d380620bfdc27901
Retrieval
retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
License
CC0 1.0 · quote-minimal attributed spans

Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents

page 3, Section 3.3

no archived snapshot exists at any timestamp
Edition
em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
Edition digest
sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
Span
em:dossier-span:sha256:73d72803976f4bb5529880a7185545e8b0ce5581409b9f64ba6ace3ac07c4e2c
Span digest
sha256:3a14d6f574bed17e76ef8b560bbca0aa27b3a50918e49461433e34ec7770a4c9
Retrieval
retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
License
CC0 1.0 · quote-minimal attributed spans

Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents

Section 3.3 'URL extraction and classification'

"URLs returning 4xx or 5xx status codes, connection errors, or timeouts are classified as non-resolving"
Edition
em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
Edition digest
sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
Span
em:dossier-span:sha256:57bdc827daee408c0042ec10327966b2b898d525449bee7ecbce1e02330f7f5f
Span digest
sha256:b240bd208cb4c322f3c9c242d01e13da53b3fd2a64e41b8b0f26653032e0ee54
Retrieval
retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
License
CC0 1.0 · quote-minimal attributed spans

Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents

Section 3 Experimental setup; Section 3.1 Datasets; section#S3.SS1; paragraph p#S3.SS1.p1.1

DRBench (17) comprises 100 multilingual research queries (Chinese and English) covering finance, science, and technology, with pre-collected outputs from 23 models.
Edition
em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
Edition digest
sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
Span
em:dossier-span:sha256:1adfd92244c349596074ef5bf78e7edf7297e8966c13f429410a0f5871d89e45
Span digest
sha256:9741316a5669b3b6e1e3022dfe4bfc23f53569a2839c46889423478a0b3c058f
Retrieval
retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
License
CC0 1.0 · quote-minimal attributed spans
07 ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.

Typed relation: support

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Limitations, lines 259-261

The benchmark primarily draws from peer-reviewed survey papers on arXiv, most of which are concentrated in STEM fields.
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:bf94ce1ab912090b1db248bca0eae62fb77c0f4f12f2fffca9cafaf6524deadd
Span digest
sha256:4947edaf643bd0a234e0ca914d920a1a5b068d18b017dcfe6d6103c466d422ca
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Section 4, line 186

the cited URL does not exist
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:e896138fd05fe0904311d35606553601c77784112c448651ce822a85e31ed99c
Span digest
sha256:ac4714fdd34ead02999923441a133a798a67254f0449f293905ef5b1d1caffab
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Section 3.3, line 155

78.87% vs. 72.94%
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:d87df9d6e418184717863f7c621ddc6c16838dc569a1ab2a558421fbc20eeaa9
Span digest
sha256:4545fdd46cab2ffc151c9ddbfaab08b82862b68eebf5205cda180b5875a3aa81
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

p. 5, lines 315-329

"we manually collected responses ... during the period from July 14 to July 25... OpenAI was using the standard version of Deep Research, powered by the o3 model."
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:3c149d37c88677f94f1715b536ad1808c4e476392f7e02212adbc723e30d8e27
Span digest
sha256:e45fbf61feb38fc8de32433d439fdb8544133e905103b884c92327930a929b0a
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Section 2.2 'Cited statements'

"retrieve the full content of each cited webpage"
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:5da64a3b244d7aa2f4baebee7a32be87f2f4255de92db8287d1b11a2a396932d
Span digest
sha256:153b2f37a8662ac0dbbd7fe51bcbb291afd5d227c4da3f4c6463274c63a11fb2
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Settings and metrics, lines 132-139

For statement extraction, supporting source extraction, and semantic consistency verification, we adopt gpt-4o.
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:f3bc53b3bd401949fea37755e5943e417ceb79e044727a73f7219cfc824ab859
Span digest
sha256:397d217cf5022ad78aa1072c2e82bcf3197e5a87cf4ef96ea81a78e1ad39b384
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Section 3.1 'Setttings'

"during the period from July 14 to July 25"
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:abc076cdbec4977f668ce36fa6e5edcfebfd81c18c90a4f7913fbb0f8d5fe33a
Span digest
sha256:b55cc3d05dbba5d160b75515ce407128fe2e2cf8c2d679517e79309d6ed6ab40
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Table 1, OpenAI Deep Research row and cited-statements Match Rate column

OpenAI Deep Research ... Match Rate 78.87%
Edition
em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
Edition digest
sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
Span
em:dossier-span:sha256:5ea28918cbd966b0fc14072426f2b7132bfe47f03d327e0b6eaee0629c3740ca
Span digest
sha256:9b1bf034e823976ac267434e9c67ee7e4038da04f3d2582a0de7d6d729fc6916
Retrieval
retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
License
CC BY 4.0 · quote-minimal attributed spans

Verify every number

Complete count ledgers

Every displayed total is a view over these typed members; no total is maintained as marketing copy.

8 · Captured reports
  1. V2-SOL-01V2-SOL-01 — completed
  2. V2-SOL-02V2-SOL-02 — completed
  3. V2-SOL-03V2-SOL-03 — completed
  4. V2-SOL-04V2-SOL-04 — completed
  5. V2-TERRA-01V2-TERRA-01 — completed
  6. V2-TERRA-02V2-TERRA-02 — completed
  7. V2-TERRA-03V2-TERRA-03 — completed
  8. V2-TERRA-04V2-TERRA-04 — completed
48 · Citation occurrences
  1. V2-SOL-01:s1_keplingerV2-SOL-01:s1_keplinger — unresolved
  2. V2-SOL-01:s1a_keplinger_dataV2-SOL-01:s1a_keplinger_data — unresolved
  3. V2-SOL-01:s2_drbenchV2-SOL-01:s2_drbench — unresolved
  4. V2-SOL-01:s2a_drbench_repoV2-SOL-01:s2a_drbench_repo — matched-exact-span
  5. V2-SOL-01:s3_deeptraceV2-SOL-01:s3_deeptrace — unresolved
  6. V2-SOL-01:s4_cited_not_verifiedV2-SOL-01:s4_cited_not_verified — unresolved
  7. V2-SOL-01:s5_url_healthV2-SOL-01:s5_url_health — unresolved
  8. V2-SOL-02:S1V2-SOL-02:S1 — unresolved
  9. V2-SOL-02:S2V2-SOL-02:S2 — matched-exact-span
  10. V2-SOL-02:S3V2-SOL-02:S3 — unresolved
  11. V2-SOL-02:S4V2-SOL-02:S4 — unresolved
  12. V2-SOL-02:S5V2-SOL-02:S5 — unresolved
  13. V2-SOL-02:S6V2-SOL-02:S6 — matched-exact-span
  14. V2-SOL-02:S7V2-SOL-02:S7 — unresolved
  15. V2-SOL-03:s1V2-SOL-03:s1 — unresolved
  16. V2-SOL-03:s2V2-SOL-03:s2 — matched-exact-span
  17. V2-SOL-03:s3V2-SOL-03:s3 — unresolved
  18. V2-SOL-03:s4V2-SOL-03:s4 — unresolved
  19. V2-SOL-03:s5V2-SOL-03:s5 — matched-exact-span
  20. V2-SOL-03:s6V2-SOL-03:s6 — matched-exact-span
  21. V2-SOL-03:s7V2-SOL-03:s7 — matched-exact-span
  22. V2-SOL-04:S1V2-SOL-04:S1 — matched-exact-span
  23. V2-SOL-04:S2V2-SOL-04:S2 — matched-exact-span
  24. V2-SOL-04:S3V2-SOL-04:S3 — matched-exact-span
  25. V2-SOL-04:S4V2-SOL-04:S4 — matched-exact-span
  26. V2-SOL-04:S5V2-SOL-04:S5 — matched-exact-span
  27. V2-SOL-04:S6V2-SOL-04:S6 — unresolved
  28. V2-SOL-04:S7V2-SOL-04:S7 — matched-exact-span
  29. V2-TERRA-01:s1_deepresearchbenchV2-TERRA-01:s1_deepresearchbench — unresolved
  30. V2-TERRA-01:s2_reportbenchV2-TERRA-01:s2_reportbench — unresolved
  31. V2-TERRA-01:s3_researcherbenchV2-TERRA-01:s3_researcherbench — unresolved
  32. V2-TERRA-01:s4_cited_not_verifiedV2-TERRA-01:s4_cited_not_verified — unresolved
  33. V2-TERRA-01:s5_urlhealthV2-TERRA-01:s5_urlhealth — unresolved
  34. V2-TERRA-02:s1_deepresearchbenchV2-TERRA-02:s1_deepresearchbench — unresolved
  35. V2-TERRA-02:s2_liveresearchbenchV2-TERRA-02:s2_liveresearchbench — unresolved
  36. V2-TERRA-02:s3_deeptraceV2-TERRA-02:s3_deeptrace — unresolved
  37. V2-TERRA-02:s4_researcherbenchV2-TERRA-02:s4_researcherbench — unresolved
  38. V2-TERRA-03:S1V2-TERRA-03:S1 — unresolved
  39. V2-TERRA-03:S2V2-TERRA-03:S2 — unresolved
  40. V2-TERRA-03:S3V2-TERRA-03:S3 — unresolved
  41. V2-TERRA-03:S4V2-TERRA-03:S4 — unresolved
  42. V2-TERRA-03:S5V2-TERRA-03:S5 — matched-exact-span
  43. V2-TERRA-03:S6V2-TERRA-03:S6 — unresolved
  44. V2-TERRA-04:S1_deepresearchbenchV2-TERRA-04:S1_deepresearchbench — unresolved
  45. V2-TERRA-04:S2_researcherbenchV2-TERRA-04:S2_researcherbench — unresolved
  46. V2-TERRA-04:S3_reportbenchV2-TERRA-04:S3_reportbench — unresolved
  47. V2-TERRA-04:S4_urlhealthV2-TERRA-04:S4_urlhealth — unresolved
  48. V2-TERRA-04:S5_cited_not_verifiedV2-TERRA-04:S5_cited_not_verified — unresolved
30 · Distinct cited URL strings
  1. https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrieved
  2. https://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrieved
  3. https://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrieved
  4. https://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrieved
  5. https://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrieved
  6. https://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrieved
  7. https://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrieved
  8. https://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrieved
  9. https://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrieved
  10. https://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrieved
  11. https://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrieved
  12. https://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrieved
  13. https://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrieved
  14. https://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrieved
  15. https://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrieved
  16. https://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrieved
  17. https://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrieved
  18. https://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrieved
  19. https://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrieved
  20. https://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
  21. https://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
  22. https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
  23. https://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrieved
  24. https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 — inaccessible
  25. https://openreview.net/pdf?id=hQ0K2Hhq7Hhttps://openreview.net/pdf?id=hQ0K2Hhq7H — inaccessible
  26. https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
  27. https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
  28. https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
  29. https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
  30. https://pubmed.ncbi.nlm.nih.gov/40904191/https://pubmed.ncbi.nlm.nih.gov/40904191/ — inaccessible
27 · Resolving URL roots
  1. https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrieved
  2. https://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrieved
  3. https://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrieved
  4. https://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrieved
  5. https://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrieved
  6. https://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrieved
  7. https://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrieved
  8. https://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrieved
  9. https://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrieved
  10. https://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrieved
  11. https://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrieved
  12. https://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrieved
  13. https://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrieved
  14. https://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrieved
  15. https://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrieved
  16. https://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrieved
  17. https://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrieved
  18. https://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrieved
  19. https://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrieved
  20. https://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
  21. https://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
  22. https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
  23. https://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrieved
  24. https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
  25. https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
  26. https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
  27. https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
11 · Source works
  1. work-citation-verifier-benchmark-2d5e94336bDo You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution — examined-source-work
  2. work-cited-not-verified-3825b25622Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — examined-source-work
  3. work-deepresearch-bench-paper-8ef36010d2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — examined-source-work
  4. work-deepresearch-bench-repository-1e475c7631Ayanami0730/deep_research_bench — examined-source-work
  5. work-deeptrace-650d66cfcbDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — examined-source-work
  6. work-keplinger-dermatology-audit-639f9ea37aAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — examined-source-work
  7. work-keplinger-supplement-1a03034fb3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — examined-source-work
  8. work-liveresearchbench-61e7da557bLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — examined-source-work
  9. work-reportbench-dca823c910ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — examined-source-work
  10. work-researcherbench-82729ab60fResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — examined-source-work
  11. work-url-health-f64342d493Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — examined-source-work
14 · Examined editions
  1. edition-citation-verifier-arxiv-v1-acfd4abab7Quote-minimal projection of edition:citation-verifier-arxiv-v1 — examined-edition
  2. edition-cnv-arxiv-v1-f24001c50aQuote-minimal projection of edition:cnv-arxiv-v1 — examined-edition
  3. edition-deeptrace-arxiv-v1-70a6e720c1Quote-minimal projection of edition:deeptrace-arxiv-v1 — examined-edition
  4. edition-deeptrace-iclr-2026-b2b153eaa6Quote-minimal projection of edition:deeptrace-iclr-2026 — examined-edition
  5. edition-drbench-arxiv-v1-7afa5ade36Quote-minimal projection of edition:drbench-arxiv-v1 — examined-edition
  6. edition-drbench-iclr-2026-d00a3dcdb6Quote-minimal projection of edition:drbench-iclr-2026 — examined-edition
  7. edition-drbench-repo-main-469cce5-4f8964f3e3Quote-minimal projection of edition:drbench-repo-main-469cce5 — examined-edition
  8. edition-keplinger-supplement-v1-f6ddb4c04aQuote-minimal projection of edition:keplinger-supplement-v1 — examined-edition
  9. edition-keplinger-supplement-v2-01e44c09d8Quote-minimal projection of edition:keplinger-supplement-v2 — examined-edition
  10. edition-keplinger-vor-2025-10ba10f860Quote-minimal projection of edition:keplinger-vor-2025 — examined-edition
  11. edition-liveresearchbench-arxiv-v5-f5644763e5Quote-minimal projection of edition:liveresearchbench-arxiv-v5 — examined-edition
  12. edition-reportbench-arxiv-v1-25888dc44bQuote-minimal projection of edition:reportbench-arxiv-v1 — examined-edition
  13. edition-researcherbench-arxiv-v1-39a8b256d7Quote-minimal projection of edition:researcherbench-arxiv-v1 — examined-edition
  14. edition-url-health-arxiv-v1-85fb934b53Quote-minimal projection of edition:url-health-arxiv-v1 — examined-edition
72 · Accepted exact span roots
  1. span-015bb68c904bf225e84c82b1acffb3cc38a4a087d7b0f79d-4c454092b1Evaluation methodology, lines 144-149 — matched-exact-span
  2. span-084f11e57cf779850ea87aaf30ba1277438ff9efcb533fd2-fe72b27cfbAppendix A.1 'Limitations' — matched-exact-span
  3. span-092745c7db6407d4d52927642741cde229fad05a5347bb18-7b04227f8fMain text, paragraph beginning 'For their high performances in generating reference lists' — matched-exact-span
  4. span-0abbab3d6269d58eb76d8e90eff1e4c27512cf46d9b8e5ff-54e5759a8brepository README, Overview — matched-exact-span
  5. span-0d4f357f2a5565bdaea01a1e537650958d71216868d1a2b0-9b0af35d76Table 1, lines 177-182 — matched-exact-span
  6. span-13cb0ddb087b82988b7445a9cfb02562f9c044ef4f244e8b-87c3b42c68Section 3.2, lines 144-146 — matched-exact-span
  7. span-196c54d6212fbc0c054dbf8e8467f2388c1f788bb93f4b1a-d460e927f6Table 1, Perplexity Deep Research citation-accuracy cell — matched-exact-span
  8. span-1db12b68cfeb4c5f5a96fa9cd2e2bd46e8dfb4f096f73154-cea47f0c65Limitations, lines 259-261 — matched-exact-span
  9. span-279db725380c133044944dccb379b75a396242219aa30437-569fc9e2fdpage 16, FACT judge validation — matched-exact-span
  10. span-2f311bfb451c6208101cb7583af37c338914fcdb10f5c75b-90e7764ce2Table 1, OpenAI Deep Research row; FACT columns — matched-exact-span
  11. span-3015a109e7898cb720e426e9a57c526277b648c75f1fb4b1-745d84ff44Section 4, line 186 — matched-exact-span
  12. span-35927851edd2e7e5e0f60d98498940f88304ba99bf1f85a0-2fe586436bpage 14, Table 3 — matched-exact-span
  13. span-39a9931f05ea356fe904ed15768fa08c66440b1fcb20f98e-191c24bbb2Table 1, ChatGPT all-correct subtotal — matched-exact-span
  14. span-417182b4c15ac550469e435a02aab1e91b395f72f186312d-e9e2c4a70cSection 4.3, line 212 — matched-exact-span
  15. span-46c392bc942cd88d525d74298e6bc6ebd9aeaec533ddfb7b-774f5b996emain text, paragraph beginning 'For their high performances'; Figure 1 — matched-exact-span
  16. span-55de6a6ca9414a252592e9f77544a9259533dd713ce1f12f-0826216ffcpp. 1-2, lines 64-95 — matched-exact-span
  17. span-59d0da64f738e1d8745f2bff3663eb89e273f93fe501afe3-ba69fc58ebSection 2 definition discussion and Section 3.3 — matched-exact-span
  18. span-5cd9ba4ca72da10cf51c0beddc6ea2765c74721bdd526a67-ed8f789629Limitations, lines 219-224 — matched-exact-span
  19. span-6885cf2524462022475b5829da081b2d35fcfe7463ca6ad0-3b1a122fafSection 1 contributions — matched-exact-span
  20. span-6c1d327e58ef001533ea1c11010e974436bbeafbde138efc-f6efd8d376Section 3.3, line 155 — matched-exact-span
  21. span-6e961498ed0f81a950f775956587f741e0eb9f11640fd3cb-afa5b59854Section 4.3, Tables 2–3 — matched-exact-span
  22. span-6f79c3c824ccf889d73e234335661e5fca68d90f8c78ee97-dbf7848a72dataset page, Description — matched-exact-span
  23. span-732c03e69c6e142a492275cee5d1e2d474c75422c9b84aaa-f7675998ccSection 3.3.1-3.3.3 — matched-exact-span
  24. span-74f6acafbb4c54d556a572460bf1b45bacd1e293f4aa16cc-5ee15c82c4Abstract, lines 47-48 — matched-exact-span
  25. span-76e9bd59b31c7ebd2ed7a2186e59753a79d3c730e0d7b4fb-9c98d99242Appendix D, Table 5 — matched-exact-span
  26. span-794bdf7e0595799afbb0d55f3dc522077aed9615c59f1e35-43f2ba944dAbstract, line 16 — matched-exact-span
  27. span-7a6046a756d6ec0315d330cf0817d062ea48498bdc9c90b3-c3a1e80029p. 5, lines 315-329 — matched-exact-span
  28. span-7d7331942f9d0520e18e96206e7718e5187788be4d48cef9-dfdbc1d019page 4, Section 3.2, Support Judgment — matched-exact-span
  29. span-886b50d94e3facabc487342cd2af5157cf956622ecfc3686-e7bdad0935Section 4.2.1, Citation Support Verification and Score Computation — matched-exact-span
  30. span-8c8d9a6dc003bf0298542b85df3f6f08db631e68e89f606f-92731379f2Table 1, ChatGPT subtotal — matched-exact-span
  31. span-8cce3d9ee1e1674ba1b93321dd9b99a55d834ca980451f8a-d2fceafe11Abstract, line 46 — matched-exact-span
  32. span-912c931063109f84ea88aea34891dd5e0f9147d2176af258-4d566ca076Abstract and Table 1 — matched-exact-span
  33. span-92eb7cf19745e24190dda84c77d590941134d09f06805c0f-88afe43377README FACT, lines 240-248 — matched-exact-span
  34. span-94ad583514aa7e04b8d5f7b8f9cdda6fdffe8ca6712f033d-63c1ba24ebDataset description — matched-exact-span
  35. span-9c7a4c186cc85f185aa293a7c5a46c08f9dbbc1a9e63c449-feff07a6e2Section 2.2 'Cited statements' — matched-exact-span
  36. span-9cb1700c08368feb234207045354cbf559d6ed80f2b0bae8-f898722eb5pp.4-7, Sections 3.1.1 and 3.2 — matched-exact-span
  37. span-9dd6eb46529e77a62748093528e52685fbd92d46a756aef7-96d75d9e2bSection 3.2, lines 163-171 — matched-exact-span
  38. span-9fe42cb48702d8d96ebbbc0309692214ece6ec495c447c06-80da541506Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment' — matched-exact-span
  39. span-a08468cb62f2feb21d9cbb95175ce00f54592bd83573cc91-03f65cfa45Official ICLR abstract — matched-exact-span
  40. span-a15ab02e9fcb23df02fb03d28967d4220ddf65ad380d2230-ddf85fa359page 4, source scraping — matched-exact-span
  41. span-a18e778fac3d49e81167f05e09fbc361e91af8c0c01b0fec-ca1b18e206page 3, Section 3.3 — matched-exact-span
  42. span-aa1534ce2e177153ab85e512deb8d064d14ec1eae16d8032-2299096a09p.14, Table 3 — matched-exact-span
  43. span-b0a39c963cdfdc5f9852cfc9f7e33e05d31232bd1d899885-adf9545222p. 4, lines 150-159 — matched-exact-span
  44. span-b38a2af5982b4c85bc205d2f533a23ed3a8f40a49b651a5f-27395a6291Main text, paragraph beginning 'To address the gap' (search-indexed primary full text) — matched-exact-span
  45. span-b68153666e4e12aee8dc887fef0a61fa5f44b5f680a0b7eb-769c2cc17dSection 5.1 Results — matched-exact-span
  46. span-b6a6ba48a2c4e352c8565b671a15a90400739f2aeefda8a6-e8232daf27Main text, following paragraph (search-indexed primary full text) — matched-exact-span
  47. span-b83996e3b5fdf878e04d6d41d0e7a1eebee9fdd9bbccf2ae-820a3eab86Appendix C, line 365 — matched-exact-span
  48. span-b903dc2284e1b4d80bb6ef948d79bbdbc7d7447e871281b8-6ef6234316body, paragraph immediately after Table 1 — matched-exact-span
  49. span-ba6b8a8b7d9121ebb05d594176ceed3d1e9dbc637876783a-d85147b6b0Settings and metrics, lines 132-139 — matched-exact-span
  50. span-bb16d5f06abe8641d9ea6694e549aa15e240dd88946aa47a-5946b6e6b7Abstract and Table 1 — matched-exact-span
  51. span-bb62928b5e0cb4e373b2f6bfeb579d8564d05da5588f8b15-771310f4efSection 3.3 'URL extraction and classification' — matched-exact-span
  52. span-c27aaca02bda06ed764be53351158fc862af9d4a556d3e82-104e7cda52page 5, Section 4.1 — matched-exact-span
  53. span-c3dfc98566c27561eb7267972e00e3e1ba6a399a3ff5d002-37f40d3260page 5, Section 3.3.3 — matched-exact-span
  54. span-c7dbd1d9e09980703ecbee8161594e4b426c1af63117b286-153d26f046Section 5 'Limitations' — matched-exact-span
  55. span-cd21c450c0d33df20a2b1a539d5725347db1b5bab7d1e986-2e3e0c3ac7Abstract, lines 45-49 — matched-exact-span
  56. span-d638c8b21f0e4a0f840880b6ffe0217ce1623758cc09fdcc-8585f6d39bSection 4.3, Tables 2-3 and following paragraph — matched-exact-span
  57. span-d880991ad784390029d49f84949d083c4789d67d45893d87-c6ef8a5c13p.33, Appendix E — matched-exact-span
  58. span-ddf4320a6d302c13d28b6225eb7e79ac15583349e1b449d4-f89bb46179Section 3.1 'Setttings' — matched-exact-span
  59. span-deada8bacb987a47f2a42e52fbfeb5c1beb3d250c0be79e2-0e91f5caccSection 4.4, line 217 — matched-exact-span
  60. span-e223640cb0ed12c87ca1d2406f3276f30a3b8d2017dd4dc1-3cfecd76bfTable 1, OpenAI Deep Research row and cited-statements Match Rate column — matched-exact-span
  61. span-e2c9f6220843a04665f5d1cd142ac15f04211ace5c579fa6-c5e0476df0Paper section 3.1.1, factual-support validation — matched-exact-span
  62. span-e4a9f30effc031388ed5c9993dc6f071903bec5e3f1998db-5d95269b76pp.17-18, Appendix E, equations 4-6 — matched-exact-span
  63. span-e54ee1457854a9a5e46f7c7c2e66ac6a5a69b963f8ee3450-8f875cfe94Abstract, line 16 — matched-exact-span
  64. span-e5bcf0241a1aeae7d9a92755b1bcf0adeeea93a8922f22f8-77a5ff4cc0README News, 2026-05-11 evaluator migration notice — matched-exact-span
  65. span-e6f447129d9a662fb202d4ac22e16822bacfdbcbcdb7afe3-f7abe2194cp.23, Appendix C, Citation Accuracy — matched-exact-span
  66. span-ea9cb2e1b0078c0857b789d61d0a0d26801cf66891864705-2f19633d83Section 5.1 and Appendix E, Table 4 — matched-exact-span
  67. span-eaf37bef7b8dc35f3e00b85d3e58ad872dc0ea604a062c8c-fa09300f85p.4, Section 3.2, lines 148-155 — matched-exact-span
  68. span-ef637b38f36195bea09205d532e8744daf3e257f93d97144-26007673a3README News, lines 176-188 — matched-exact-span
  69. span-f05b3189dc3fafd45b5bde643a91fd0616832928215a37a2-47f72e77e4Paragraph after Table 1 — matched-exact-span
  70. span-f35c9565f2fd047f6e7272dc608bbf6a63813a717bad5354-e6a07f05b6Appendix C, line 365 — matched-exact-span
  71. span-f96ad5daa81153a3689f52f5777fedfd70e8a87d44e7a1c9-3da72b8748Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3) — matched-exact-span
  72. span-f98f1274859ba98cc538e3be6d11d345a5382cfcb9516aa1-d60b765c48p.17, Appendix C, lines 704-708 — matched-exact-span
7 · Candidate warrant roots
  1. relation-cnv-link-relevance-support-gap-165443d270Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.
  2. relation-cnv-search-depth-ablation-9c01321092Within the Cited but Not Verified harness, the 2-to-150-call ablation reduced reported Fact Check scores for two setups without establishing a general causal law about search depth.
  3. relation-deepresearch-bench-fact-a3f61d7a3aDeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.
  4. relation-deeptrace-support-variation-1a84ef7c6fDeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.
  5. relation-keplinger-metadata-versus-support-c23e105c8bIn one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.
  6. relation-reportbench-match-rate-6964e41a69ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.
  7. relation-url-health-resolution-0b26fb45c7The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.
0 · Independently confirmed warrant roots
    4 · Pending warrant groups
    1. relation-citation-verifier-calibration-5b40082faaThe captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    2. relation-liveresearchbench-e1-e2-e3-5fb5d671f7The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    3. relation-researcherbench-faithfulness-groundedness-c84dcaee54The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    4. relation-url-health-correction-loop-eb8a6384c6The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
    34 · Unresolved citation occurrences
    1. V2-SOL-01:s1_keplingerAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    2. V2-SOL-01:s1a_keplinger_dataSupplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    3. V2-SOL-01:s2_drbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    4. V2-SOL-01:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    5. V2-SOL-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    6. V2-SOL-01:s5_url_healthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    7. V2-SOL-02:S1Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    8. V2-SOL-02:S3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    9. V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    10. V2-SOL-02:S5ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    11. V2-SOL-02:S7Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    12. V2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    13. V2-SOL-03:s3DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    14. V2-SOL-03:s4Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    15. V2-SOL-04:S6Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
    16. V2-TERRA-01:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    17. V2-TERRA-01:s2_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    18. V2-TERRA-01:s3_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    19. V2-TERRA-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    20. V2-TERRA-01:s5_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    21. V2-TERRA-02:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    22. V2-TERRA-02:s2_liveresearchbenchLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — unresolved
    23. V2-TERRA-02:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
    24. V2-TERRA-02:s4_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    25. V2-TERRA-03:S1Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    26. V2-TERRA-03:S2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    27. V2-TERRA-03:S3ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    28. V2-TERRA-03:S4Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    29. V2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    30. V2-TERRA-04:S1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    31. V2-TERRA-04:S2_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
    32. V2-TERRA-04:S3_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
    33. V2-TERRA-04:S4_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
    34. V2-TERRA-04:S5_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
    20 · Unsupported or force-raised claims
    1. V2-SOL-01:r3_deepresearch_bench_factV2-SOL-01:r3_deepresearch_bench_fact — no-credit
    2. V2-SOL-01:r4_deeptrace_supportV2-SOL-01:r4_deeptrace_support — no-credit
    3. V2-SOL-02:R3_deepresearch_bench_factV2-SOL-02:R3_deepresearch_bench_fact — no-credit
    4. V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — no-credit
    5. V2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — no-credit
    6. V2-SOL-03:r4_deeptrace_supportV2-SOL-03:r4_deeptrace_support — no-credit
    7. V2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — no-credit
    8. V2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — no-credit
    9. V2-SOL-04:R1V2-SOL-04:R1 — no-credit
    10. V2-SOL-04:R2V2-SOL-04:R2 — no-credit
    11. V2-SOL-04:R3V2-SOL-04:R3 — no-credit
    12. V2-SOL-04:R4V2-SOL-04:R4 — no-credit
    13. V2-SOL-04:R5V2-SOL-04:R5 — no-credit
    14. V2-SOL-04:R8V2-SOL-04:R8 — no-credit
    15. V2-TERRA-01:r1_deepresearchbench_factV2-TERRA-01:r1_deepresearchbench_fact — no-credit
    16. V2-TERRA-01:r4_cited_not_verified_source_attributionV2-TERRA-01:r4_cited_not_verified_source_attribution — no-credit
    17. V2-TERRA-02:answerV2-TERRA-02:answer — no-credit
    18. V2-TERRA-02:r1_deepresearchbench_factV2-TERRA-02:r1_deepresearchbench_fact — no-credit
    19. V2-TERRA-02:r4_deeptraceV2-TERRA-02:r4_deeptrace — no-credit
    20. V2-TERRA-03:R3V2-TERRA-03:R3 — no-credit
    9 · Independently rejected claims
    1. V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — independently-rejected
    2. V2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — independently-rejected
    3. V2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — independently-rejected
    4. V2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — independently-rejected
    5. V2-SOL-04:R1V2-SOL-04:R1 — independently-rejected
    6. V2-SOL-04:R2V2-SOL-04:R2 — independently-rejected
    7. V2-SOL-04:R3V2-SOL-04:R3 — independently-rejected
    8. V2-SOL-04:R4V2-SOL-04:R4 — independently-rejected
    9. V2-SOL-04:R8V2-SOL-04:R8 — independently-rejected
    3 · Inaccessible carriers
    1. V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
    2. V2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
    3. V2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved

    Five-line lexicon

    The units are different on purpose

    Citation occurrence
    One report-level use of a citation, before repeated URLs are collapsed.
    URL root
    One distinct cited URL string; resolution does not prove claim support.
    Source work
    One logical paper, dataset, repository, or audit instrument across editions.
    Candidate warrant root
    One source-method-data proposition that survived bounded semantic review, while residual independence remains unresolved.
    No credit
    The item remains visible but does not support the stronger claim because its carrier, span, semantics, or lineage did not close.

    Build receipt

    Reproduce this projection

    Reproducible projection
    Dossier
    em:dossier:sha256:cbd7a14096a956f642f5c76046d3b49ed648fbe6bf24144c992404a01415af82
    View policy
    em:application-policy:encyclopedia-v0.1
    Catalog
    em:catalog:sha256:092898e1fe3d355761ab4cec653576926a8f5d31621ec7ce23dd60e9d19563ef
    Frontier
    em:frontier:sha256:7e4a173112ef26422acf3ed9434c8b6849c4e011797e20fed6c0a9ca58a1e4c3
    Accepted commit
    af081caa99fc08d3fabb914ff68f2e672a83bd5b
    Epistemic policy
    commons-balanced-v0.1
    Disclosure policy
    public-noninterference-v0.1
    Compiler
    epistemedia/0.2.0
    Content digest
    048c12622d9daca7cd009a7483c58697aa0674f77b12c6eca3721954ec1f3743