How We Know · Case 002 · encyclopedia
When eight research agents agree, how many evidence roots are there?
What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?
Eight context-isolated public-by-design reports captured under one frozen prompt, plus the public source editions and exact spans they cited. This historical pilot does not estimate current or universal agent behavior.
encyclopedia finding
The bounded record contains empirical methods for URL resolution and claim-to-source support, but the eight reports reuse overlapping capture, source, method, and derivation lineages. Report agreement alone adds no independent warrant beyond the inspected source record.
What to do with that: Use a polished cited report as a map into sources, not as a vote count. Check URL resolution and sentence-level support separately.
This bounded 2026 packet does not estimate current or universal agent reliability.
Lineage accounting
Agreement is not a vote count
Shared capture: Unknown provider and retrieval dependencies remain; all reports share the exact prompt and one bounded capture program, so run multiplicity gets zero automatic independence credit.
Warrant boundary: Unknown residual independence remains across task data, judge methods, retrieval, source, edition, span, derivation, and upstream-citation lineages.
Sentence x-ray
What the selected record actually supports
Open a sentence to inspect its exact work, edition, span, retrieval, digest, and license chain.
01 This report is a captured observation, not an independent evidence root.
Typed relation: dependence
EM-0026 deterministic agent-citation evidence ledger
Captured report V2-SOL-01
{ "answer_bytes": 40476, "answer_path": "research/how-we-know/agent-citation-lineage/answers-v2/V2-SOL-01.json", "answer_sha256": "17a7cf97b717d7b520026b9eb52a7da59ddb9b5f698c2aeadfab3f38c91008a0", "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_model_profile": "gpt-5.6-sol", "retrieval_infrastructure": "unknown", "run_id": "V2-SOL-01", "status": "completed", "trace_path": "research/how-we-know/agent-citation-lineage/traces-v2/V2-SOL-01.json" }
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:ead238d33280cfd1dd252936d47fc1351950f8efb4105085d161ac7c1e51fa97
- Span digest
- sha256:ab05cfa54280c980dabd52b0d336f1dc65a5761925a54dad2d5c7b1cbfb094cb
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
EM-0026 deterministic agent-citation evidence ledger
Shared capture lineage
{ "automatic_independence_credit": 0, "prompt_sha256": "d321a9cec7b5fe419157c0623e18ff0020cb0080fbe6dd9ed4fd20a0b896f670", "reported_model_identity": "unknown", "requested_profiles": [ "gpt-5.6-sol", "gpt-5.6-terra" ], "retrieval_infrastructure": "unknown" }
- Edition
- em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0
- Edition digest
- sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b
- Span
- em:dossier-span:sha256:e120064fee0bdabe15299cb865ff2f7df66473291a13b7a765a4a09d16223f5e
- Span digest
- sha256:92b558b23559a96a551de8acb2627b148688d21ede2ca3fb9966196c4675dab6
- Retrieval
- No external retrieval record; repository audit span
- License
- Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments
02 In one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.
Typed relation: support
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Main text, paragraph beginning 'For their high performances in generating reference lists'
citation-bearing sentences exhibited high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:9f91508aa7be57f8e16c902e5399763ce349e7318b10d6085b65f0f53d5bdb56
- Span digest
- sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Table 1, ChatGPT all-correct subtotal
16 (69.6%)
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:72a3edb3976aa397d38e9a7ccbc18b5e4af6f00c91540cd47aadcf9eddce2439
- Span digest
- sha256:b6c3ff6f2ff60f37ae1c6b773ba6aa5a6c01015b61e3cc6f2a06bfcef69b53ad
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
main text, paragraph beginning 'For their high performances'; Figure 1
citation-bearing sentences exhibited high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:910f7758be3883fc0964cd1e5412b9c75a586e6c4c7c08e64ebda5b9f541a132
- Span digest
- sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Table 1, ChatGPT subtotal
23 (100.0%)
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:d9dab2e8e8da55f5a5ddbe3e62c658a7f643cd05c36c117aa8a0dfc76a7b6d09
- Span digest
- sha256:5428bd1a04d5d2e332015107e915fe3a38f1d5e173e8c7246b743d98ecfd8618
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype
Dataset description
sentence-level evaluation
- Edition
- em:dossier-edition:sha256:8c80a50cd40cbb31bdaedb6730e8e173e3edc0f3440538dbcc4878eb0997ffd9
- Edition digest
- sha256:0a485e13dcc56e5f883213fde7e24bdcc4ffb810bda8ac95186f83f6325fafd1
- Span
- em:dossier-span:sha256:38bf64d13c4edef9c0aba131c84d42e0e08637a661d2552cbf2cb49536813488
- Span digest
- sha256:e3ec5c81f615c657f7c6b124c5b2150d86f18281384ca0bd24ebb43a3317796f
- Retrieval
- retrieved https://data.mendeley.com/datasets/3s73z9zf3c/1
- License
- CC BY 4.0 · metadata and quote-minimal landing-page spans
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Main text, paragraph beginning 'To address the gap' (search-indexed primary full text)
We imposed a 1500-word limit, APA-style references and conducted three independent runs.
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:54d8710861e7e1ee03c50697f1316c12138246e303a08c843fdadecfa99b8c5e
- Span digest
- sha256:1c94be957cfe750f445dc9143722905b5784b02f754a9b4071a3381e2bd493e9
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Main text, following paragraph (search-indexed primary full text)
ChatGPT Deep Research-generated references were mostly identifiable by their metadata (95.7%): 69.6% were entirely correct
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:b9e4b58f255c32e2a9e5e6796e8d35f7f28d7481060d6b35f1e02a3c74d2b3d6
- Span digest
- sha256:28141fd1b0222658eb93a338783358926de770223eea24f8c5cbcf632d67f955
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
body, paragraph immediately after Table 1
high error rates: 51.3 ± 6.5% for ChatGPT and 57.8 ± 22.7% for Le Chat
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:f687ae7c3889980d4d3eea1b63b7bea25ae674ebc04caf8c59e7ae4f4fb77464
- Span digest
- sha256:2e43a0652065a6a1fb69de224082ac79d903cc71cac0664a1cb376cbfa09d0e8
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype
Paragraph after Table 1
51.3 ± 6.5%
- Edition
- em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761
- Edition digest
- sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097
- Span
- em:dossier-span:sha256:2f331f1e2636907b0cab1000b0947e440fa5f24eec0edaf9498470522e1abe1a
- Span digest
- sha256:1f863c1f35c37af1fd2c3e4927ea0b289d94a9532fa5855e209ce98ea344d498
- Retrieval
- retrieved https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ · inaccessible https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 · inaccessible https://pubmed.ncbi.nlm.nih.gov/40904191/
- License
- CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution
03 DeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.
Typed relation: support
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Evaluation methodology, lines 144-149
Each unique Statement-URL pair undergoes a support evaluation.
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:ed08815f4aabd455caf37f61a1e37a2dc35725098f2e9be5c6da6c9d5854d612
- Span digest
- sha256:1bde849b0f5df91e847e907dd763f27be01ffdc7eb0c7e61898c8a9731472c35
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Table 1, lines 177-182
Perplexity Deep Research | ... | 90.24 | 31.26 ... Gemini-2.5-Pro Deep Research | ... | 81.44 | 111.21 ... OpenAI Deep Research | ... | 77.96 | 40.79
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:fd8835d2aaa9358bd7dfbcf07848ed19001fa82f11e45ceff3374cdd4e882be3
- Span digest
- sha256:53d38f9cf0d0aedf28a77d7de32935372202fff78a0c02d98f8c269ef223906a
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Section 3.2, lines 144-146
‘support’ or ‘not support’
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:4648142e8019a2d54f201844685fbb87a921d93fc24b0f5c08f8124a82ffc374
- Span digest
- sha256:a49c200ea10833496ac73719992be09d5fd3fd021fdd51ccc587ac2592225f56
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Table 1, Perplexity Deep Research citation-accuracy cell
90.24
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:b118f44725ec505886f92a94123fe289be763523b8afb669757c72ccd0d8c352
- Span digest
- sha256:64c911242980826f92ffae6d0571646ecc3aceca73c213791c6d3a8d889b4cc7
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
page 16, FACT judge validation
aligned with human 'support' determinations in 96% of cases
- Edition
- em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
- Edition digest
- sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
- Span
- em:dossier-span:sha256:a59ed2a9ea0e6972bc3f71d63feb1c361b76eb8d8fa5bb679c7beab61a32b713
- Span digest
- sha256:8be6257e74c3696813fb0079c27e93e8bec16c71d214b6da43811fbd9b406ba4
- Retrieval
- inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
- License
- CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Table 1, OpenAI Deep Research row; FACT columns
OpenAI Deep Research | 46.98 | 46.87 | 45.25 | 49.27 | 47.14 | 77.96 | 40.79
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:762ebe0afa7a2ddf8eb19d879c91f63d1de7d25192f10d248b65fefb0e113038
- Span digest
- sha256:dedd156999fa7d57ae9c22a166c56c263547eee3c37e6f41fe7644d980110986
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
page 4, Section 3.2, Support Judgment
This yields a binary judgment: 'support' or 'not support'.
- Edition
- em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
- Edition digest
- sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
- Span
- em:dossier-span:sha256:d915faef70b54c6751577ba10f4057744b2b1b0b4d10613cdd8d74720a5da23b
- Span digest
- sha256:78f021c07980218ead835a986c84780444479e627dcfa9f2ea126b3604b89ad7
- Retrieval
- inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
- License
- CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only
Ayanami0730/deep_research_bench
README FACT, lines 240-248
Support Verification: Uses web scraping and LLM judgment to verify whether cited sources actually support the claims
- Edition
- em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
- Edition digest
- sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
- Span
- em:dossier-span:sha256:a28651fed7ffef6fdbbe7f322faa2482aa1e0f6b1e2a0a909eddef59a72740f2
- Span digest
- sha256:549018c6df261f27bd10133ad4e4f565e004f07710095259a7889cc79d149794
- Retrieval
- retrieved https://github.com/Ayanami0730/deep_research_bench
- License
- Apache-2.0 · metadata and quote-minimal README spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment'
"Each unique Statement-URL pair undergoes a support evaluation."
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:aba61c5b1685e20da29c0d6a42dd49b66f0b6c2f3ca08483716ae5c518562d4f
- Span digest
- sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
p. 4, lines 150-159
"Each unique Statement-URL pair undergoes a support evaluation... This results in a binary judgment ('support' or 'not support') for each pair, determining whether the citation accurately grounds the claim."
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:0b565d8281beb5c9b95d300108446049bc4c93de4dc3e80a501f41e85d0a98a5
- Span digest
- sha256:4edb75ec3f690eb2d15089923ba5e9c4e6174f55094d0919041bf10625901ceb
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Appendix C, line 365
96% of cases
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:d753bc9c860da325190141f3f78a0d031f2d6c7cf5e1a51f19f355f44cfb47db
- Span digest
- sha256:ae845f85955fa47d52d51f32803e6a26afa8b725c523e53d9ef30b2abca72d9b
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
page 5, Section 4.1
complete set of 100 tasks
- Edition
- em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1
- Edition digest
- sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f
- Span
- em:dossier-span:sha256:4aa9441f93251fa969f1fe8cc8a4fb6cf728ec3f8879efae3ffff55fabae759d
- Span digest
- sha256:11034cf2238003f213bf87210c9017f57cc743910fa70754f3e61f875abe75e6
- Retrieval
- inaccessible https://openreview.net/pdf?id=hQ0K2Hhq7H · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf
- License
- CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
pp.17-18, Appendix E, equations 4-6
"the proportion of 'support' statement-URL pairs for each individual task"
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:c6710bb4b411353c7c600e449cc556e39a3530dc63c002a7b3973ef6bb1793bf
- Span digest
- sha256:82b9361b6e696a92b41fdfae500f56b05d48575c0d429f15b3a8d6786c83b8d7
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
Ayanami0730/deep_research_bench
README News, 2026-05-11 evaluator migration notice
Legacy code: the previous Gemini-2.5-Pro / Gemini-2.5-Flash evaluation code is preserved on the Gemini-2.5 branch.
- Edition
- em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
- Edition digest
- sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
- Span
- em:dossier-span:sha256:abc4ab811e5c574575a0024557dda57a45d8d439ad4ff4ce08bef3753f7a86ef
- Span digest
- sha256:02cd1c04d6312fcda76549885095d9c9d971cc14709ae8f749cf47fc2c11bcee
- Retrieval
- retrieved https://github.com/Ayanami0730/deep_research_bench
- License
- Apache-2.0 · metadata and quote-minimal README spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
p.4, Section 3.2, lines 148-155
"Each unique Statement-URL pair undergoes a support evaluation."
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:27d667cb8cff21b646ec759c6c6fe79c5744a28079337d7a7fa9ac98abf8da9d
- Span digest
- sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
Ayanami0730/deep_research_bench
README News, lines 176-188
Official Evaluator Switched to GPT-5.5 ... Legacy code: the previous Gemini-2.5-Pro / Gemini-2.5-Flash evaluation code is preserved
- Edition
- em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b
- Edition digest
- sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b
- Span
- em:dossier-span:sha256:fd761963efd1f957ebe8dae1eb76239c2d17ce1553109a4e0ac420525c73cfcb
- Span digest
- sha256:9c713be4fc872b46e66d812e033708b51b4e385c81aa45f7b1798459a677343d
- Retrieval
- retrieved https://github.com/Ayanami0730/deep_research_bench
- License
- Apache-2.0 · metadata and quote-minimal README spans
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
p.17, Appendix C, lines 704-708
"aligned with human 'support' determinations in 96% of cases"
- Edition
- em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c
- Edition digest
- sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c
- Span
- em:dossier-span:sha256:17193f2aab7106eece041bcabda963d9bd172d798792efc058b6a93c9bd4cb0c
- Span digest
- sha256:43c892b582d42a335c982009f97347b5bf2ae2799cc3e478c53504f4d01e4388
- Retrieval
- retrieved https://arxiv.org/html/2506.11763 · retrieved https://arxiv.org/html/2506.11763v1 · retrieved https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf
- License
- CC BY 4.0; unknown · quote-minimal attributed spans
04 DeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.
Typed relation: support
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
page 14, Table 3
Factual support (statement-source) 0.62 binary
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:6cf2883b02dd404844a662bae3367ab60d91f6847160a3df9c6ec6b287dc0858
- Span digest
- sha256:bb4f28cd0f4527c633bcfacf51b18ced7387c4a2f266c59d41de7ee899b24b85
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Abstract, line 46
citation accuracy ranging from 40–80% across systems
- Edition
- em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b
- Edition digest
- sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d
- Span
- em:dossier-span:sha256:e748230bbc2f6b31433eaba4536917834c454cb73a7f18e8d9b975e2cac447b5
- Span digest
- sha256:a0eeca5c86f9c7aeda9167f626a8cb0ed56dc87ea1ac5ca8c4a7a0add76e7625
- Retrieval
- retrieved https://arxiv.org/html/2509.04499 · retrieved https://arxiv.org/html/2509.04499v1
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
pp.4-7, Sections 3.1.1 and 3.2
"The dataset comprises 303 questions"
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:b9cbe6b1d35e74f3c8133d2be41a47d9361a26958e9f644dafd97e9e5acb516a
- Span digest
- sha256:2ce91759797b86455a478a9ced58ec59d032e1a32a1c964347f96be4ab9e2054
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Section 3.2, lines 163-171
303 queries x 9 models
- Edition
- em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b
- Edition digest
- sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d
- Span
- em:dossier-span:sha256:8c1ac40e63a414c1a41cd0842632677269c8800967f22ef8247b6fb36f6242ea
- Span digest
- sha256:5cc83ab699cb1342475b407e34d33fc106d64b5d9270e5bb9445d80586ebabc2
- Retrieval
- retrieved https://arxiv.org/html/2509.04499 · retrieved https://arxiv.org/html/2509.04499v1
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Official ICLR abstract
with citation accuracy ranging from 40–80% across systems
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:06ea026aae360f700369f4b1f658c4ca85e8acce431923cd7978ff2358b1ec05
- Span digest
- sha256:af76864a50c641f8d693216d3b7ebe2a414eebf75467305c1e1a6946b6746590
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
page 4, source scraping
For roughly 15% of the URLs, the Reader tool returns an error
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:12ab62355d178256bbaed6cdf5cbf6ae150410d9364f98cced347b54161fa5e8
- Span digest
- sha256:a38c9ac275570f2b78dd1e1bf7663579e35f9037131e12d721dfb4c6f005683d
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
p.14, Table 3
"Factual support (statement–source) 0.62"
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:2aa386b3653a99db3b670f783272056371480ed287dc1f62a51d3d587d680996
- Span digest
- sha256:75a9212dc448c15ffc1e4a36bdf190728977470661bd9f45be534010704faa25
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Paper section 3.1.1, factual-support validation
Pearson correlation of 0.62 between the LLM judge and manual labels
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:a575506a3cca26acda376aa922fdb0e269767f71b01f502a7c50f3d4492e6f03
- Span digest
- sha256:d2f6a4561a938ee913dadde2fc4a80c32c492fbcb3b47f236c9639375843ae8d
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
printed page 10; PDF page index 10/22; Section 4 Results; Deep Research Agents paragraph
Gemini(DR) demonstrates weak citation performance: only 40.3% citation accuracy
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:533550e262b8691ce9ed33aa0e89a7bf13d89c04c6002e1855ed1e767f556e3a
- Span digest
- sha256:cf32a686cec566152749e3a4c643fa29aa115099f9eba77c0ca113c48264f769
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
printed page 9; PDF page index 9/22; Table 1; Gemini (DR) column; %Citation Accuracy row
{ "cell": "50.3", "column": "Gemini (DR)", "row": "%Citation Accuracy" }
- Edition
- em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70
- Edition digest
- sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545
- Span
- em:dossier-span:sha256:cf4e063e6f297fe901b342752b67850d84cc3e8e21ace2ef8d5ea4f3fd93091c
- Span digest
- sha256:e020a09063eee7aa1eb70663cf3e8bd7a8162952f5ae65d959d57ef74ab855f6
- Retrieval
- retrieved https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html · retrieved https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf
- License
- arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only
05 Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.
Typed relation: support
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Section 3.3.1-3.3.3
"Fact Check verifies whether specific factual claims are accurately supported by the source content."
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:d4e2f35392b064365ad867f33ead8636bb8bbf3af03c0141585e5ae56c868090
- Span digest
- sha256:a4386d56325583893d806fdd15ccd475d5096f0b445b1b9f52aeba253c910987
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Abstract, lines 47-48
yet achieve only 39–77% factual accuracy
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:ea33bc17365ed01ed68ab06f33c4fd2020dc8e6947659e4e9629b44af4aef635
- Span digest
- sha256:be05510d58cc0231bf968f29c70b3d669a82d5cd0c74a778a5f8c4e035bf45b4
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Abstract and Table 1
link validity above 94% and relevance above 80%, yet achieve only 39–77% factual accuracy
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:49ee66722d6d523dd485f05c796181af55736c3ceab2a235c5c3b40cca95e439
- Span digest
- sha256:44d4c6941f2920e907de1e4cd22aeb9a725f24cc19432723754e316bd6f5cc49
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Abstract and Table 1
maintain link validity above 94% and relevance above 80%, yet achieve only 39–77% factual accuracy
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:8cf2ff4c32b2dac6d6dfc7684dd2bf7dca9d40fe7b7b36d50dfdcddad067bff3
- Span digest
- sha256:e35ca66d12349d7290a6dbe7449bf6da173eb0671765644b581713528b316c74
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Abstract, lines 45-49
Citations are evaluated along three dimensions. (1) Link Works verifies URL accessibility, (2) Relevant Content measures topical alignment, and (3) Fact Check validates factual accuracy against source content.
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:237508b682de825711f61762deb5ce7661d5d197e9fde3aed5f5e80f05fdcdb5
- Span digest
- sha256:d34096a9d9b1fbb43ef50ba503c3894607c008b563e1e1a5d1643cbc23c5d447
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
Section 4.4, line 217
only 1 failed link out of 2,159 evaluations
- Edition
- em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29
- Edition digest
- sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b
- Span
- em:dossier-span:sha256:2bd67a5caa606894f4301ca1185c88a4dde02d3c5c4a2cf311463a0f39332443
- Span digest
- sha256:da34a8e2a1d4e190d3b2095d79886b3cbf7edb2c63874189aa381302d2c2924e
- Retrieval
- retrieved https://arxiv.org/html/2605.06635 · retrieved https://arxiv.org/html/2605.06635v1 · retrieved https://arxiv.org/pdf/2605.06635 · retrieved https://arxiv.org/abs/2605.06635
- License
- CC BY 4.0 · quote-minimal attributed spans
06 The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.
Typed relation: support
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
pp. 1-2, lines 64-95
"This study focuses on URL-based citation hallucinations"; "fabricated snippets ... and invented bibliographic entries ... require separate systematic study."
- Edition
- em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
- Edition digest
- sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
- Span
- em:dossier-span:sha256:179697ec62048a6465fa50dda34ff54530e700430013a878201b9dadde405cc0
- Span digest
- sha256:ee1de7374c67e208aa5d1331c20a08c8c83418309f1bd768d380620bfdc27901
- Retrieval
- retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
- License
- CC0 1.0 · quote-minimal attributed spans
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
page 3, Section 3.3
no archived snapshot exists at any timestamp
- Edition
- em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
- Edition digest
- sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
- Span
- em:dossier-span:sha256:73d72803976f4bb5529880a7185545e8b0ce5581409b9f64ba6ace3ac07c4e2c
- Span digest
- sha256:3a14d6f574bed17e76ef8b560bbca0aa27b3a50918e49461433e34ec7770a4c9
- Retrieval
- retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
- License
- CC0 1.0 · quote-minimal attributed spans
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
Section 3.3 'URL extraction and classification'
"URLs returning 4xx or 5xx status codes, connection errors, or timeouts are classified as non-resolving"
- Edition
- em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
- Edition digest
- sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
- Span
- em:dossier-span:sha256:57bdc827daee408c0042ec10327966b2b898d525449bee7ecbce1e02330f7f5f
- Span digest
- sha256:b240bd208cb4c322f3c9c242d01e13da53b3fd2a64e41b8b0f26653032e0ee54
- Retrieval
- retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
- License
- CC0 1.0 · quote-minimal attributed spans
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
Section 3 Experimental setup; Section 3.1 Datasets; section#S3.SS1; paragraph p#S3.SS1.p1.1
DRBench (17) comprises 100 multilingual research queries (Chinese and English) covering finance, science, and technology, with pre-collected outputs from 23 models.
- Edition
- em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f
- Edition digest
- sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7
- Span
- em:dossier-span:sha256:1adfd92244c349596074ef5bf78e7edf7297e8966c13f429410a0f5871d89e45
- Span digest
- sha256:9741316a5669b3b6e1e3022dfe4bfc23f53569a2839c46889423478a0b3c058f
- Retrieval
- retrieved https://arxiv.org/html/2604.03173 · retrieved https://arxiv.org/html/2604.03173v1 · retrieved https://arxiv.org/pdf/2604.03173 · retrieved https://arxiv.org/abs/2604.03173
- License
- CC0 1.0 · quote-minimal attributed spans
07 ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.
Typed relation: support
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Limitations, lines 259-261
The benchmark primarily draws from peer-reviewed survey papers on arXiv, most of which are concentrated in STEM fields.
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:bf94ce1ab912090b1db248bca0eae62fb77c0f4f12f2fffca9cafaf6524deadd
- Span digest
- sha256:4947edaf643bd0a234e0ca914d920a1a5b068d18b017dcfe6d6103c466d422ca
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Section 4, line 186
the cited URL does not exist
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:e896138fd05fe0904311d35606553601c77784112c448651ce822a85e31ed99c
- Span digest
- sha256:ac4714fdd34ead02999923441a133a798a67254f0449f293905ef5b1d1caffab
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Section 3.3, line 155
78.87% vs. 72.94%
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:d87df9d6e418184717863f7c621ddc6c16838dc569a1ab2a558421fbc20eeaa9
- Span digest
- sha256:4545fdd46cab2ffc151c9ddbfaab08b82862b68eebf5205cda180b5875a3aa81
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
p. 5, lines 315-329
"we manually collected responses ... during the period from July 14 to July 25... OpenAI was using the standard version of Deep Research, powered by the o3 model."
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:3c149d37c88677f94f1715b536ad1808c4e476392f7e02212adbc723e30d8e27
- Span digest
- sha256:e45fbf61feb38fc8de32433d439fdb8544133e905103b884c92327930a929b0a
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Section 2.2 'Cited statements'
"retrieve the full content of each cited webpage"
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:5da64a3b244d7aa2f4baebee7a32be87f2f4255de92db8287d1b11a2a396932d
- Span digest
- sha256:153b2f37a8662ac0dbbd7fe51bcbb291afd5d227c4da3f4c6463274c63a11fb2
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Settings and metrics, lines 132-139
For statement extraction, supporting source extraction, and semantic consistency verification, we adopt gpt-4o.
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:f3bc53b3bd401949fea37755e5943e417ceb79e044727a73f7219cfc824ab859
- Span digest
- sha256:397d217cf5022ad78aa1072c2e82bcf3197e5a87cf4ef96ea81a78e1ad39b384
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Section 3.1 'Setttings'
"during the period from July 14 to July 25"
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:abc076cdbec4977f668ce36fa6e5edcfebfd81c18c90a4f7913fbb0f8d5fe33a
- Span digest
- sha256:b55cc3d05dbba5d160b75515ce407128fe2e2cf8c2d679517e79309d6ed6ab40
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Table 1, OpenAI Deep Research row and cited-statements Match Rate column
OpenAI Deep Research ... Match Rate 78.87%
- Edition
- em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6
- Edition digest
- sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8
- Span
- em:dossier-span:sha256:5ea28918cbd966b0fc14072426f2b7132bfe47f03d327e0b6eaee0629c3740ca
- Span digest
- sha256:9b1bf034e823976ac267434e9c67ee7e4038da04f3d2582a0de7d6d729fc6916
- Retrieval
- retrieved https://arxiv.org/html/2508.15804 · retrieved https://arxiv.org/html/2508.15804v1 · retrieved https://arxiv.org/pdf/2508.15804
- License
- CC BY 4.0 · quote-minimal attributed spans
Verify every number
Complete count ledgers
Every displayed total is a view over these typed members; no total is maintained as marketing copy.
8 · Captured reports
V2-SOL-01V2-SOL-01 — completedV2-SOL-02V2-SOL-02 — completedV2-SOL-03V2-SOL-03 — completedV2-SOL-04V2-SOL-04 — completedV2-TERRA-01V2-TERRA-01 — completedV2-TERRA-02V2-TERRA-02 — completedV2-TERRA-03V2-TERRA-03 — completedV2-TERRA-04V2-TERRA-04 — completed
48 · Citation occurrences
V2-SOL-01:s1_keplingerV2-SOL-01:s1_keplinger — unresolvedV2-SOL-01:s1a_keplinger_dataV2-SOL-01:s1a_keplinger_data — unresolvedV2-SOL-01:s2_drbenchV2-SOL-01:s2_drbench — unresolvedV2-SOL-01:s2a_drbench_repoV2-SOL-01:s2a_drbench_repo — matched-exact-spanV2-SOL-01:s3_deeptraceV2-SOL-01:s3_deeptrace — unresolvedV2-SOL-01:s4_cited_not_verifiedV2-SOL-01:s4_cited_not_verified — unresolvedV2-SOL-01:s5_url_healthV2-SOL-01:s5_url_health — unresolvedV2-SOL-02:S1V2-SOL-02:S1 — unresolvedV2-SOL-02:S2V2-SOL-02:S2 — matched-exact-spanV2-SOL-02:S3V2-SOL-02:S3 — unresolvedV2-SOL-02:S4V2-SOL-02:S4 — unresolvedV2-SOL-02:S5V2-SOL-02:S5 — unresolvedV2-SOL-02:S6V2-SOL-02:S6 — matched-exact-spanV2-SOL-02:S7V2-SOL-02:S7 — unresolvedV2-SOL-03:s1V2-SOL-03:s1 — unresolvedV2-SOL-03:s2V2-SOL-03:s2 — matched-exact-spanV2-SOL-03:s3V2-SOL-03:s3 — unresolvedV2-SOL-03:s4V2-SOL-03:s4 — unresolvedV2-SOL-03:s5V2-SOL-03:s5 — matched-exact-spanV2-SOL-03:s6V2-SOL-03:s6 — matched-exact-spanV2-SOL-03:s7V2-SOL-03:s7 — matched-exact-spanV2-SOL-04:S1V2-SOL-04:S1 — matched-exact-spanV2-SOL-04:S2V2-SOL-04:S2 — matched-exact-spanV2-SOL-04:S3V2-SOL-04:S3 — matched-exact-spanV2-SOL-04:S4V2-SOL-04:S4 — matched-exact-spanV2-SOL-04:S5V2-SOL-04:S5 — matched-exact-spanV2-SOL-04:S6V2-SOL-04:S6 — unresolvedV2-SOL-04:S7V2-SOL-04:S7 — matched-exact-spanV2-TERRA-01:s1_deepresearchbenchV2-TERRA-01:s1_deepresearchbench — unresolvedV2-TERRA-01:s2_reportbenchV2-TERRA-01:s2_reportbench — unresolvedV2-TERRA-01:s3_researcherbenchV2-TERRA-01:s3_researcherbench — unresolvedV2-TERRA-01:s4_cited_not_verifiedV2-TERRA-01:s4_cited_not_verified — unresolvedV2-TERRA-01:s5_urlhealthV2-TERRA-01:s5_urlhealth — unresolvedV2-TERRA-02:s1_deepresearchbenchV2-TERRA-02:s1_deepresearchbench — unresolvedV2-TERRA-02:s2_liveresearchbenchV2-TERRA-02:s2_liveresearchbench — unresolvedV2-TERRA-02:s3_deeptraceV2-TERRA-02:s3_deeptrace — unresolvedV2-TERRA-02:s4_researcherbenchV2-TERRA-02:s4_researcherbench — unresolvedV2-TERRA-03:S1V2-TERRA-03:S1 — unresolvedV2-TERRA-03:S2V2-TERRA-03:S2 — unresolvedV2-TERRA-03:S3V2-TERRA-03:S3 — unresolvedV2-TERRA-03:S4V2-TERRA-03:S4 — unresolvedV2-TERRA-03:S5V2-TERRA-03:S5 — matched-exact-spanV2-TERRA-03:S6V2-TERRA-03:S6 — unresolvedV2-TERRA-04:S1_deepresearchbenchV2-TERRA-04:S1_deepresearchbench — unresolvedV2-TERRA-04:S2_researcherbenchV2-TERRA-04:S2_researcherbench — unresolvedV2-TERRA-04:S3_reportbenchV2-TERRA-04:S3_reportbench — unresolvedV2-TERRA-04:S4_urlhealthV2-TERRA-04:S4_urlhealth — unresolvedV2-TERRA-04:S5_cited_not_verifiedV2-TERRA-04:S5_cited_not_verified — unresolved
30 · Distinct cited URL strings
https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrievedhttps://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrievedhttps://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrievedhttps://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrievedhttps://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrievedhttps://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrievedhttps://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrievedhttps://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrievedhttps://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrievedhttps://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrievedhttps://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrievedhttps://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrievedhttps://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrievedhttps://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrievedhttps://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrievedhttps://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrievedhttps://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrievedhttps://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrievedhttps://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrievedhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrievedhttps://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrievedhttps://onlinelibrary.wiley.com/doi/10.1111/jdv.70035https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 — inaccessiblehttps://openreview.net/pdf?id=hQ0K2Hhq7Hhttps://openreview.net/pdf?id=hQ0K2Hhq7H — inaccessiblehttps://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrievedhttps://pubmed.ncbi.nlm.nih.gov/40904191/https://pubmed.ncbi.nlm.nih.gov/40904191/ — inaccessible
27 · Resolving URL roots
https://arxiv.org/abs/2604.03173https://arxiv.org/abs/2604.03173 — retrievedhttps://arxiv.org/abs/2605.06635https://arxiv.org/abs/2605.06635 — retrievedhttps://arxiv.org/abs/2607.08700https://arxiv.org/abs/2607.08700 — retrievedhttps://arxiv.org/html/2506.11763https://arxiv.org/html/2506.11763 — retrievedhttps://arxiv.org/html/2506.11763v1https://arxiv.org/html/2506.11763v1 — retrievedhttps://arxiv.org/html/2507.16280https://arxiv.org/html/2507.16280 — retrievedhttps://arxiv.org/html/2508.15804https://arxiv.org/html/2508.15804 — retrievedhttps://arxiv.org/html/2508.15804v1https://arxiv.org/html/2508.15804v1 — retrievedhttps://arxiv.org/html/2509.04499https://arxiv.org/html/2509.04499 — retrievedhttps://arxiv.org/html/2509.04499v1https://arxiv.org/html/2509.04499v1 — retrievedhttps://arxiv.org/html/2604.03173https://arxiv.org/html/2604.03173 — retrievedhttps://arxiv.org/html/2604.03173v1https://arxiv.org/html/2604.03173v1 — retrievedhttps://arxiv.org/html/2605.06635https://arxiv.org/html/2605.06635 — retrievedhttps://arxiv.org/html/2605.06635v1https://arxiv.org/html/2605.06635v1 — retrievedhttps://arxiv.org/pdf/2507.16280https://arxiv.org/pdf/2507.16280 — retrievedhttps://arxiv.org/pdf/2508.15804https://arxiv.org/pdf/2508.15804 — retrievedhttps://arxiv.org/pdf/2510.14240v5https://arxiv.org/pdf/2510.14240v5 — retrievedhttps://arxiv.org/pdf/2604.03173https://arxiv.org/pdf/2604.03173 — retrievedhttps://arxiv.org/pdf/2605.06635https://arxiv.org/pdf/2605.06635 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/1https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrievedhttps://data.mendeley.com/datasets/3s73z9zf3c/2https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrievedhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdfhttps://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrievedhttps://github.com/Ayanami0730/deep_research_benchhttps://github.com/Ayanami0730/deep_research_bench — retrievedhttps://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdfhttps://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrievedhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.htmlhttps://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
11 · Source works
work-citation-verifier-benchmark-2d5e94336bDo You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution — examined-source-workwork-cited-not-verified-3825b25622Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — examined-source-workwork-deepresearch-bench-paper-8ef36010d2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — examined-source-workwork-deepresearch-bench-repository-1e475c7631Ayanami0730/deep_research_bench — examined-source-workwork-deeptrace-650d66cfcbDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — examined-source-workwork-keplinger-dermatology-audit-639f9ea37aAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — examined-source-workwork-keplinger-supplement-1a03034fb3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — examined-source-workwork-liveresearchbench-61e7da557bLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — examined-source-workwork-reportbench-dca823c910ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — examined-source-workwork-researcherbench-82729ab60fResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — examined-source-workwork-url-health-f64342d493Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — examined-source-work
14 · Examined editions
edition-citation-verifier-arxiv-v1-acfd4abab7Quote-minimal projection of edition:citation-verifier-arxiv-v1 — examined-editionedition-cnv-arxiv-v1-f24001c50aQuote-minimal projection of edition:cnv-arxiv-v1 — examined-editionedition-deeptrace-arxiv-v1-70a6e720c1Quote-minimal projection of edition:deeptrace-arxiv-v1 — examined-editionedition-deeptrace-iclr-2026-b2b153eaa6Quote-minimal projection of edition:deeptrace-iclr-2026 — examined-editionedition-drbench-arxiv-v1-7afa5ade36Quote-minimal projection of edition:drbench-arxiv-v1 — examined-editionedition-drbench-iclr-2026-d00a3dcdb6Quote-minimal projection of edition:drbench-iclr-2026 — examined-editionedition-drbench-repo-main-469cce5-4f8964f3e3Quote-minimal projection of edition:drbench-repo-main-469cce5 — examined-editionedition-keplinger-supplement-v1-f6ddb4c04aQuote-minimal projection of edition:keplinger-supplement-v1 — examined-editionedition-keplinger-supplement-v2-01e44c09d8Quote-minimal projection of edition:keplinger-supplement-v2 — examined-editionedition-keplinger-vor-2025-10ba10f860Quote-minimal projection of edition:keplinger-vor-2025 — examined-editionedition-liveresearchbench-arxiv-v5-f5644763e5Quote-minimal projection of edition:liveresearchbench-arxiv-v5 — examined-editionedition-reportbench-arxiv-v1-25888dc44bQuote-minimal projection of edition:reportbench-arxiv-v1 — examined-editionedition-researcherbench-arxiv-v1-39a8b256d7Quote-minimal projection of edition:researcherbench-arxiv-v1 — examined-editionedition-url-health-arxiv-v1-85fb934b53Quote-minimal projection of edition:url-health-arxiv-v1 — examined-edition
72 · Accepted exact span roots
span-015bb68c904bf225e84c82b1acffb3cc38a4a087d7b0f79d-4c454092b1Evaluation methodology, lines 144-149 — matched-exact-spanspan-084f11e57cf779850ea87aaf30ba1277438ff9efcb533fd2-fe72b27cfbAppendix A.1 'Limitations' — matched-exact-spanspan-092745c7db6407d4d52927642741cde229fad05a5347bb18-7b04227f8fMain text, paragraph beginning 'For their high performances in generating reference lists' — matched-exact-spanspan-0abbab3d6269d58eb76d8e90eff1e4c27512cf46d9b8e5ff-54e5759a8brepository README, Overview — matched-exact-spanspan-0d4f357f2a5565bdaea01a1e537650958d71216868d1a2b0-9b0af35d76Table 1, lines 177-182 — matched-exact-spanspan-13cb0ddb087b82988b7445a9cfb02562f9c044ef4f244e8b-87c3b42c68Section 3.2, lines 144-146 — matched-exact-spanspan-196c54d6212fbc0c054dbf8e8467f2388c1f788bb93f4b1a-d460e927f6Table 1, Perplexity Deep Research citation-accuracy cell — matched-exact-spanspan-1db12b68cfeb4c5f5a96fa9cd2e2bd46e8dfb4f096f73154-cea47f0c65Limitations, lines 259-261 — matched-exact-spanspan-279db725380c133044944dccb379b75a396242219aa30437-569fc9e2fdpage 16, FACT judge validation — matched-exact-spanspan-2f311bfb451c6208101cb7583af37c338914fcdb10f5c75b-90e7764ce2Table 1, OpenAI Deep Research row; FACT columns — matched-exact-spanspan-3015a109e7898cb720e426e9a57c526277b648c75f1fb4b1-745d84ff44Section 4, line 186 — matched-exact-spanspan-35927851edd2e7e5e0f60d98498940f88304ba99bf1f85a0-2fe586436bpage 14, Table 3 — matched-exact-spanspan-39a9931f05ea356fe904ed15768fa08c66440b1fcb20f98e-191c24bbb2Table 1, ChatGPT all-correct subtotal — matched-exact-spanspan-417182b4c15ac550469e435a02aab1e91b395f72f186312d-e9e2c4a70cSection 4.3, line 212 — matched-exact-spanspan-46c392bc942cd88d525d74298e6bc6ebd9aeaec533ddfb7b-774f5b996emain text, paragraph beginning 'For their high performances'; Figure 1 — matched-exact-spanspan-55de6a6ca9414a252592e9f77544a9259533dd713ce1f12f-0826216ffcpp. 1-2, lines 64-95 — matched-exact-spanspan-59d0da64f738e1d8745f2bff3663eb89e273f93fe501afe3-ba69fc58ebSection 2 definition discussion and Section 3.3 — matched-exact-spanspan-5cd9ba4ca72da10cf51c0beddc6ea2765c74721bdd526a67-ed8f789629Limitations, lines 219-224 — matched-exact-spanspan-6885cf2524462022475b5829da081b2d35fcfe7463ca6ad0-3b1a122fafSection 1 contributions — matched-exact-spanspan-6c1d327e58ef001533ea1c11010e974436bbeafbde138efc-f6efd8d376Section 3.3, line 155 — matched-exact-spanspan-6e961498ed0f81a950f775956587f741e0eb9f11640fd3cb-afa5b59854Section 4.3, Tables 2–3 — matched-exact-spanspan-6f79c3c824ccf889d73e234335661e5fca68d90f8c78ee97-dbf7848a72dataset page, Description — matched-exact-spanspan-732c03e69c6e142a492275cee5d1e2d474c75422c9b84aaa-f7675998ccSection 3.3.1-3.3.3 — matched-exact-spanspan-74f6acafbb4c54d556a572460bf1b45bacd1e293f4aa16cc-5ee15c82c4Abstract, lines 47-48 — matched-exact-spanspan-76e9bd59b31c7ebd2ed7a2186e59753a79d3c730e0d7b4fb-9c98d99242Appendix D, Table 5 — matched-exact-spanspan-794bdf7e0595799afbb0d55f3dc522077aed9615c59f1e35-43f2ba944dAbstract, line 16 — matched-exact-spanspan-7a6046a756d6ec0315d330cf0817d062ea48498bdc9c90b3-c3a1e80029p. 5, lines 315-329 — matched-exact-spanspan-7d7331942f9d0520e18e96206e7718e5187788be4d48cef9-dfdbc1d019page 4, Section 3.2, Support Judgment — matched-exact-spanspan-886b50d94e3facabc487342cd2af5157cf956622ecfc3686-e7bdad0935Section 4.2.1, Citation Support Verification and Score Computation — matched-exact-spanspan-8c8d9a6dc003bf0298542b85df3f6f08db631e68e89f606f-92731379f2Table 1, ChatGPT subtotal — matched-exact-spanspan-8cce3d9ee1e1674ba1b93321dd9b99a55d834ca980451f8a-d2fceafe11Abstract, line 46 — matched-exact-spanspan-912c931063109f84ea88aea34891dd5e0f9147d2176af258-4d566ca076Abstract and Table 1 — matched-exact-spanspan-92eb7cf19745e24190dda84c77d590941134d09f06805c0f-88afe43377README FACT, lines 240-248 — matched-exact-spanspan-94ad583514aa7e04b8d5f7b8f9cdda6fdffe8ca6712f033d-63c1ba24ebDataset description — matched-exact-spanspan-9c7a4c186cc85f185aa293a7c5a46c08f9dbbc1a9e63c449-feff07a6e2Section 2.2 'Cited statements' — matched-exact-spanspan-9cb1700c08368feb234207045354cbf559d6ed80f2b0bae8-f898722eb5pp.4-7, Sections 3.1.1 and 3.2 — matched-exact-spanspan-9dd6eb46529e77a62748093528e52685fbd92d46a756aef7-96d75d9e2bSection 3.2, lines 163-171 — matched-exact-spanspan-9fe42cb48702d8d96ebbbc0309692214ece6ec495c447c06-80da541506Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment' — matched-exact-spanspan-a08468cb62f2feb21d9cbb95175ce00f54592bd83573cc91-03f65cfa45Official ICLR abstract — matched-exact-spanspan-a15ab02e9fcb23df02fb03d28967d4220ddf65ad380d2230-ddf85fa359page 4, source scraping — matched-exact-spanspan-a18e778fac3d49e81167f05e09fbc361e91af8c0c01b0fec-ca1b18e206page 3, Section 3.3 — matched-exact-spanspan-aa1534ce2e177153ab85e512deb8d064d14ec1eae16d8032-2299096a09p.14, Table 3 — matched-exact-spanspan-b0a39c963cdfdc5f9852cfc9f7e33e05d31232bd1d899885-adf9545222p. 4, lines 150-159 — matched-exact-spanspan-b38a2af5982b4c85bc205d2f533a23ed3a8f40a49b651a5f-27395a6291Main text, paragraph beginning 'To address the gap' (search-indexed primary full text) — matched-exact-spanspan-b68153666e4e12aee8dc887fef0a61fa5f44b5f680a0b7eb-769c2cc17dSection 5.1 Results — matched-exact-spanspan-b6a6ba48a2c4e352c8565b671a15a90400739f2aeefda8a6-e8232daf27Main text, following paragraph (search-indexed primary full text) — matched-exact-spanspan-b83996e3b5fdf878e04d6d41d0e7a1eebee9fdd9bbccf2ae-820a3eab86Appendix C, line 365 — matched-exact-spanspan-b903dc2284e1b4d80bb6ef948d79bbdbc7d7447e871281b8-6ef6234316body, paragraph immediately after Table 1 — matched-exact-spanspan-ba6b8a8b7d9121ebb05d594176ceed3d1e9dbc637876783a-d85147b6b0Settings and metrics, lines 132-139 — matched-exact-spanspan-bb16d5f06abe8641d9ea6694e549aa15e240dd88946aa47a-5946b6e6b7Abstract and Table 1 — matched-exact-spanspan-bb62928b5e0cb4e373b2f6bfeb579d8564d05da5588f8b15-771310f4efSection 3.3 'URL extraction and classification' — matched-exact-spanspan-c27aaca02bda06ed764be53351158fc862af9d4a556d3e82-104e7cda52page 5, Section 4.1 — matched-exact-spanspan-c3dfc98566c27561eb7267972e00e3e1ba6a399a3ff5d002-37f40d3260page 5, Section 3.3.3 — matched-exact-spanspan-c7dbd1d9e09980703ecbee8161594e4b426c1af63117b286-153d26f046Section 5 'Limitations' — matched-exact-spanspan-cd21c450c0d33df20a2b1a539d5725347db1b5bab7d1e986-2e3e0c3ac7Abstract, lines 45-49 — matched-exact-spanspan-d638c8b21f0e4a0f840880b6ffe0217ce1623758cc09fdcc-8585f6d39bSection 4.3, Tables 2-3 and following paragraph — matched-exact-spanspan-d880991ad784390029d49f84949d083c4789d67d45893d87-c6ef8a5c13p.33, Appendix E — matched-exact-spanspan-ddf4320a6d302c13d28b6225eb7e79ac15583349e1b449d4-f89bb46179Section 3.1 'Setttings' — matched-exact-spanspan-deada8bacb987a47f2a42e52fbfeb5c1beb3d250c0be79e2-0e91f5caccSection 4.4, line 217 — matched-exact-spanspan-e223640cb0ed12c87ca1d2406f3276f30a3b8d2017dd4dc1-3cfecd76bfTable 1, OpenAI Deep Research row and cited-statements Match Rate column — matched-exact-spanspan-e2c9f6220843a04665f5d1cd142ac15f04211ace5c579fa6-c5e0476df0Paper section 3.1.1, factual-support validation — matched-exact-spanspan-e4a9f30effc031388ed5c9993dc6f071903bec5e3f1998db-5d95269b76pp.17-18, Appendix E, equations 4-6 — matched-exact-spanspan-e54ee1457854a9a5e46f7c7c2e66ac6a5a69b963f8ee3450-8f875cfe94Abstract, line 16 — matched-exact-spanspan-e5bcf0241a1aeae7d9a92755b1bcf0adeeea93a8922f22f8-77a5ff4cc0README News, 2026-05-11 evaluator migration notice — matched-exact-spanspan-e6f447129d9a662fb202d4ac22e16822bacfdbcbcdb7afe3-f7abe2194cp.23, Appendix C, Citation Accuracy — matched-exact-spanspan-ea9cb2e1b0078c0857b789d61d0a0d26801cf66891864705-2f19633d83Section 5.1 and Appendix E, Table 4 — matched-exact-spanspan-eaf37bef7b8dc35f3e00b85d3e58ad872dc0ea604a062c8c-fa09300f85p.4, Section 3.2, lines 148-155 — matched-exact-spanspan-ef637b38f36195bea09205d532e8744daf3e257f93d97144-26007673a3README News, lines 176-188 — matched-exact-spanspan-f05b3189dc3fafd45b5bde643a91fd0616832928215a37a2-47f72e77e4Paragraph after Table 1 — matched-exact-spanspan-f35c9565f2fd047f6e7272dc608bbf6a63813a717bad5354-e6a07f05b6Appendix C, line 365 — matched-exact-spanspan-f96ad5daa81153a3689f52f5777fedfd70e8a87d44e7a1c9-3da72b8748Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3) — matched-exact-spanspan-f98f1274859ba98cc538e3be6d11d345a5382cfcb9516aa1-d60b765c48p.17, Appendix C, lines 704-708 — matched-exact-span
7 · Candidate warrant roots
relation-cnv-link-relevance-support-gap-165443d270Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.relation-cnv-search-depth-ablation-9c01321092Within the Cited but Not Verified harness, the 2-to-150-call ablation reduced reported Fact Check scores for two setups without establishing a general causal law about search depth.relation-deepresearch-bench-fact-a3f61d7a3aDeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.relation-deeptrace-support-variation-1a84ef7c6fDeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.relation-keplinger-metadata-versus-support-c23e105c8bIn one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.relation-reportbench-match-rate-6964e41a69ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.relation-url-health-resolution-0b26fb45c7The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.
0 · Independently confirmed warrant roots
4 · Pending warrant groups
relation-citation-verifier-calibration-5b40082faaThe captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-liveresearchbench-e1-e2-e3-5fb5d671f7The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-researcherbench-faithfulness-groundedness-c84dcaee54The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.relation-url-health-correction-loop-eb8a6384c6The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
34 · Unresolved citation occurrences
V2-SOL-01:s1_keplingerAssessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-SOL-01:s1a_keplinger_dataSupplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-SOL-01:s2_drbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-01:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-SOL-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-SOL-01:s5_url_healthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-02:S1Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-02:S3Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-02:S5ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-SOL-02:S7Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-SOL-03:s3DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-SOL-03:s4Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-SOL-04:S6Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolvedV2-TERRA-01:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-01:s2_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-01:s3_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-01:s4_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-TERRA-01:s5_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-TERRA-02:s1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-02:s2_liveresearchbenchLiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — unresolvedV2-TERRA-02:s3_deeptraceDeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolvedV2-TERRA-02:s4_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-03:S1Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolvedV2-TERRA-03:S2DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-03:S3ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-03:S4Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-04:S1_deepresearchbenchDeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-TERRA-04:S2_researcherbenchResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolvedV2-TERRA-04:S3_reportbenchReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolvedV2-TERRA-04:S4_urlhealthDetecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolvedV2-TERRA-04:S5_cited_not_verifiedCited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
20 · Unsupported or force-raised claims
V2-SOL-01:r3_deepresearch_bench_factV2-SOL-01:r3_deepresearch_bench_fact — no-creditV2-SOL-01:r4_deeptrace_supportV2-SOL-01:r4_deeptrace_support — no-creditV2-SOL-02:R3_deepresearch_bench_factV2-SOL-02:R3_deepresearch_bench_fact — no-creditV2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — no-creditV2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — no-creditV2-SOL-03:r4_deeptrace_supportV2-SOL-03:r4_deeptrace_support — no-creditV2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — no-creditV2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — no-creditV2-SOL-04:R1V2-SOL-04:R1 — no-creditV2-SOL-04:R2V2-SOL-04:R2 — no-creditV2-SOL-04:R3V2-SOL-04:R3 — no-creditV2-SOL-04:R4V2-SOL-04:R4 — no-creditV2-SOL-04:R5V2-SOL-04:R5 — no-creditV2-SOL-04:R8V2-SOL-04:R8 — no-creditV2-TERRA-01:r1_deepresearchbench_factV2-TERRA-01:r1_deepresearchbench_fact — no-creditV2-TERRA-01:r4_cited_not_verified_source_attributionV2-TERRA-01:r4_cited_not_verified_source_attribution — no-creditV2-TERRA-02:answerV2-TERRA-02:answer — no-creditV2-TERRA-02:r1_deepresearchbench_factV2-TERRA-02:r1_deepresearchbench_fact — no-creditV2-TERRA-02:r4_deeptraceV2-TERRA-02:r4_deeptrace — no-creditV2-TERRA-03:R3V2-TERRA-03:R3 — no-credit
9 · Independently rejected claims
V2-SOL-02:R5_deeptrace_auditV2-SOL-02:R5_deeptrace_audit — independently-rejectedV2-SOL-03:r3_deepresearch_bench_factV2-SOL-03:r3_deepresearch_bench_fact — independently-rejectedV2-SOL-03:r8_search_depth_ablationV2-SOL-03:r8_search_depth_ablation — independently-rejectedV2-SOL-03:r9_verifier_calibrationV2-SOL-03:r9_verifier_calibration — independently-rejectedV2-SOL-04:R1V2-SOL-04:R1 — independently-rejectedV2-SOL-04:R2V2-SOL-04:R2 — independently-rejectedV2-SOL-04:R3V2-SOL-04:R3 — independently-rejectedV2-SOL-04:R4V2-SOL-04:R4 — independently-rejectedV2-SOL-04:R8V2-SOL-04:R8 — independently-rejected
3 · Inaccessible carriers
V2-SOL-02:S4DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolvedV2-SOL-03:s1Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolvedV2-TERRA-03:S6PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
Five-line lexicon
The units are different on purpose
- Citation occurrence
- One report-level use of a citation, before repeated URLs are collapsed.
- URL root
- One distinct cited URL string; resolution does not prove claim support.
- Source work
- One logical paper, dataset, repository, or audit instrument across editions.
- Candidate warrant root
- One source-method-data proposition that survived bounded semantic review, while residual independence remains unresolved.
- No credit
- The item remains visible but does not support the stronger claim because its carrier, span, semantics, or lineage did not close.