# Case 002 — When eight research agents agree, how many evidence roots are there?

What empirical evidence published or publicly posted by 2026-08-22 measures whether citations produced by deep-research agents resolve and actually support the claims made from them?

**Encyclopedia finding:** The bounded record contains empirical methods for URL resolution and claim-to-source support, but the eight reports reuse overlapping capture, source, method, and derivation lineages. Report agreement alone adds no independent warrant beyond the inspected source record.

**Practical reading:** Use a polished cited report as a map into sources, not as a vote count. Check URL resolution and sentence-level support separately.

This bounded 2026 packet does not estimate current or universal agent reliability.

## Evidence accounting

- [8 captured reports](#reports): Observations from one frozen capture program
- [30 distinct URL strings](#cited-urls): A resolving link is not sentence support
- [11 source works](#source-works): Logical works after edition collapse
- [7 candidate warrants](#candidate-warrants): Scoped candidates, not independent programs
- [34 unresolved citations](#unresolved-citations): Visible and assigned no credit

**Dependence warning:** Unknown provider and retrieval dependencies remain; all reports share the exact prompt and one bounded capture program, so run multiplicity gets zero automatic independence credit.

## What the selected record says

### Dependence

This report is a captured observation, not an independent evidence root.

- Source: [EM-0026 deterministic agent-citation evidence ledger](https://github.com/yoheinakajima/epistemedia/tree/main/research/how-we-know/agent-citation-lineage)
- Work: `em:dossier-source-work:sha256:c552ac33696c6615477f161b374123595dd33681366befc990cbee17731e6e8b`
- Edition: `em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0` · `sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b`
- Span: `em:dossier-span:sha256:ead238d33280cfd1dd252936d47fc1351950f8efb4105085d161ac7c1e51fa97` · `sha256:ab05cfa54280c980dabd52b0d336f1dc65a5761925a54dad2d5c7b1cbfb094cb`
- Locator: Captured report V2-SOL-01
- License: Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments

- Source: [EM-0026 deterministic agent-citation evidence ledger](https://github.com/yoheinakajima/epistemedia/tree/main/research/how-we-know/agent-citation-lineage)
- Work: `em:dossier-source-work:sha256:c552ac33696c6615477f161b374123595dd33681366befc990cbee17731e6e8b`
- Edition: `em:dossier-edition:sha256:f0111e28ede7d9e17c435927cba483d0d9a7acf3dc23b38322ee4677d8cb70c0` · `sha256:8400ea3bfe6c1e2516ab92710486eff3e62f5fee783a38859b383a561181533b`
- Span: `em:dossier-span:sha256:e120064fee0bdabe15299cb865ff2f7df66473291a13b7a765a4a09d16223f5e` · `sha256:92b558b23559a96a551de8acb2627b148688d21ede2ca3fb9966196c4675dab6`
- Locator: Shared capture lineage
- License: Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments · Repository metadata and derived audit relations under Apache-2.0; embedded source excerpts retain their recorded treatments

### Support

In one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:9f91508aa7be57f8e16c902e5399763ce349e7318b10d6085b65f0f53d5bdb56` · `sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b`
- Locator: Main text, paragraph beginning 'For their high performances in generating reference lists'
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:72a3edb3976aa397d38e9a7ccbc18b5e4af6f00c91540cd47aadcf9eddce2439` · `sha256:b6c3ff6f2ff60f37ae1c6b773ba6aa5a6c01015b61e3cc6f2a06bfcef69b53ad`
- Locator: Table 1, ChatGPT all-correct subtotal
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:910f7758be3883fc0964cd1e5412b9c75a586e6c4c7c08e64ebda5b9f541a132` · `sha256:39f799b4bf1bcc50a74e5b019c585b970fa01779cf4b07bdeb27d71f35ea455b`
- Locator: main text, paragraph beginning 'For their high performances'; Figure 1
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:d9dab2e8e8da55f5a5ddbe3e62c658a7f643cd05c36c117aa8a0dfc76a7b6d09` · `sha256:5428bd1a04d5d2e332015107e915fe3a38f1d5e173e8c7246b743d98ecfd8618`
- Locator: Table 1, ChatGPT subtotal
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype](https://data.mendeley.com/datasets/3s73z9zf3c/1)
- Work: `em:dossier-source-work:sha256:daa2447b3107ce9818b1c435de19c6ea8ca0de07129761714bda29dc7a062eb4`
- Edition: `em:dossier-edition:sha256:8c80a50cd40cbb31bdaedb6730e8e173e3edc0f3440538dbcc4878eb0997ffd9` · `sha256:0a485e13dcc56e5f883213fde7e24bdcc4ffb810bda8ac95186f83f6325fafd1`
- Span: `em:dossier-span:sha256:38bf64d13c4edef9c0aba131c84d42e0e08637a661d2552cbf2cb49536813488` · `sha256:e3ec5c81f615c657f7c6b124c5b2150d86f18281384ca0bd24ebb43a3317796f`
- Locator: Dataset description
- License: CC BY 4.0 · metadata and quote-minimal landing-page spans

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:54d8710861e7e1ee03c50697f1316c12138246e303a08c843fdadecfa99b8c5e` · `sha256:1c94be957cfe750f445dc9143722905b5784b02f754a9b4071a3381e2bd493e9`
- Locator: Main text, paragraph beginning 'To address the gap' (search-indexed primary full text)
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:b9e4b58f255c32e2a9e5e6796e8d35f7f28d7481060d6b35f1e02a3c74d2b3d6` · `sha256:28141fd1b0222658eb93a338783358926de770223eea24f8c5cbcf632d67f955`
- Locator: Main text, following paragraph (search-indexed primary full text)
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:f687ae7c3889980d4d3eea1b63b7bea25ae674ebc04caf8c59e7ae4f4fb77464` · `sha256:2e43a0652065a6a1fb69de224082ac79d903cc71cac0664a1cb376cbfa09d0e8`
- Locator: body, paragraph immediately after Table 1
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

- Source: [Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype](https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/)
- Work: `em:dossier-source-work:sha256:7e292f6ce412d254983ada9a5032ef5fac626876b16404d99ed1a99ce4e3ba05`
- Edition: `em:dossier-edition:sha256:ecbbc5366e2c53ebf6358f0f4c06b7076e158210edd3a081f16af7ab0b029761` · `sha256:98431c96e776a0502a6387574529817629ca50e164d8cd4e7f6997215ec7e097`
- Span: `em:dossier-span:sha256:2f331f1e2636907b0cab1000b0947e440fa5f24eec0edaf9498470522e1abe1a` · `sha256:1f863c1f35c37af1fd2c3e4927ea0b289d94a9532fa5855e209ce98ea344d498`
- Locator: Paragraph after Table 1
- License: CC BY-NC 4.0 · quote-minimal attributed spans; no full-text redistribution

### Support

DeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:ed08815f4aabd455caf37f61a1e37a2dc35725098f2e9be5c6da6c9d5854d612` · `sha256:1bde849b0f5df91e847e907dd763f27be01ffdc7eb0c7e61898c8a9731472c35`
- Locator: Evaluation methodology, lines 144-149
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:fd8835d2aaa9358bd7dfbcf07848ed19001fa82f11e45ceff3374cdd4e882be3` · `sha256:53d38f9cf0d0aedf28a77d7de32935372202fff78a0c02d98f8c269ef223906a`
- Locator: Table 1, lines 177-182
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:4648142e8019a2d54f201844685fbb87a921d93fc24b0f5c08f8124a82ffc374` · `sha256:a49c200ea10833496ac73719992be09d5fd3fd021fdd51ccc587ac2592225f56`
- Locator: Section 3.2, lines 144-146
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:b118f44725ec505886f92a94123fe289be763523b8afb669757c72ccd0d8c352` · `sha256:64c911242980826f92ffae6d0571646ecc3aceca73c213791c6d3a8d889b4cc7`
- Locator: Table 1, Perplexity Deep Research citation-accuracy cell
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1` · `sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f`
- Span: `em:dossier-span:sha256:a59ed2a9ea0e6972bc3f71d63feb1c361b76eb8d8fa5bb679c7beab61a32b713` · `sha256:8be6257e74c3696813fb0079c27e93e8bec16c71d214b6da43811fbd9b406ba4`
- Locator: page 16, FACT judge validation
- License: CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:762ebe0afa7a2ddf8eb19d879c91f63d1de7d25192f10d248b65fefb0e113038` · `sha256:dedd156999fa7d57ae9c22a166c56c263547eee3c37e6f41fe7644d980110986`
- Locator: Table 1, OpenAI Deep Research row; FACT columns
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1` · `sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f`
- Span: `em:dossier-span:sha256:d915faef70b54c6751577ba10f4057744b2b1b0b4d10613cdd8d74720a5da23b` · `sha256:78f021c07980218ead835a986c84780444479e627dcfa9f2ea126b3604b89ad7`
- Locator: page 4, Section 3.2, Support Judgment
- License: CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [Ayanami0730/deep_research_bench](https://github.com/Ayanami0730/deep_research_bench/commit/469cce54ea7f6a63c163d3d9fec879cf289ec484)
- Work: `em:dossier-source-work:sha256:433604e28e4a1ee90dfc7f2c5b2389d04e64a600930ad51cd9acf599b4df9f21`
- Edition: `em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b` · `sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b`
- Span: `em:dossier-span:sha256:a28651fed7ffef6fdbbe7f322faa2482aa1e0f6b1e2a0a909eddef59a72740f2` · `sha256:549018c6df261f27bd10133ad4e4f565e004f07710095259a7889cc79d149794`
- Locator: README FACT, lines 240-248
- License: Apache-2.0 · metadata and quote-minimal README spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:aba61c5b1685e20da29c0d6a42dd49b66f0b6c2f3ca08483716ae5c518562d4f` · `sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce`
- Locator: Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment'
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:0b565d8281beb5c9b95d300108446049bc4c93de4dc3e80a501f41e85d0a98a5` · `sha256:4edb75ec3f690eb2d15089923ba5e9c4e6174f55094d0919041bf10625901ceb`
- Locator: p. 4, lines 150-159
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:d753bc9c860da325190141f3f78a0d031f2d6c7cf5e1a51f19f355f44cfb47db` · `sha256:ae845f85955fa47d52d51f32803e6a26afa8b725c523e53d9ef30b2abca72d9b`
- Locator: Appendix C, line 365
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:47d8d33326a498097021dc5754365cbb8f5abfd3e628e90908d2342f0b6c7cf1` · `sha256:02f53bf5ae17ab2bffc25ba8dd5ca912d35b85ae6ae78a202cd79fe7f7efe25f`
- Span: `em:dossier-span:sha256:4aa9441f93251fa969f1fe8cc8a4fb6cf728ec3f8879efae3ffff55fabae759d` · `sha256:11034cf2238003f213bf87210c9017f57cc743910fa70754f3e61f875abe75e6`
- Locator: page 5, Section 4.1
- License: CC BY 4.0; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:c6710bb4b411353c7c600e449cc556e39a3530dc63c002a7b3973ef6bb1793bf` · `sha256:82b9361b6e696a92b41fdfae500f56b05d48575c0d429f15b3a8d6786c83b8d7`
- Locator: pp.17-18, Appendix E, equations 4-6
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [Ayanami0730/deep_research_bench](https://github.com/Ayanami0730/deep_research_bench/commit/469cce54ea7f6a63c163d3d9fec879cf289ec484)
- Work: `em:dossier-source-work:sha256:433604e28e4a1ee90dfc7f2c5b2389d04e64a600930ad51cd9acf599b4df9f21`
- Edition: `em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b` · `sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b`
- Span: `em:dossier-span:sha256:abc4ab811e5c574575a0024557dda57a45d8d439ad4ff4ce08bef3753f7a86ef` · `sha256:02cd1c04d6312fcda76549885095d9c9d971cc14709ae8f749cf47fc2c11bcee`
- Locator: README News, 2026-05-11 evaluator migration notice
- License: Apache-2.0 · metadata and quote-minimal README spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:27d667cb8cff21b646ec759c6c6fe79c5744a28079337d7a7fa9ac98abf8da9d` · `sha256:3cca518ed14d7451346a35623b20e7f9e350e4cb5cbcec3d7f622cc6aac509ce`
- Locator: p.4, Section 3.2, lines 148-155
- License: CC BY 4.0; unknown · quote-minimal attributed spans

- Source: [Ayanami0730/deep_research_bench](https://github.com/Ayanami0730/deep_research_bench/commit/469cce54ea7f6a63c163d3d9fec879cf289ec484)
- Work: `em:dossier-source-work:sha256:433604e28e4a1ee90dfc7f2c5b2389d04e64a600930ad51cd9acf599b4df9f21`
- Edition: `em:dossier-edition:sha256:bf5c65b640b321fb75bf8b7f96854f151f834a7d31a2fc9d2bfb783852e7262b` · `sha256:35350075b3514f668c0deddb5bd4850ec31c7bfda2cf1e363995d68c580fe81b`
- Span: `em:dossier-span:sha256:fd761963efd1f957ebe8dae1eb76239c2d17ce1553109a4e0ac420525c73cfcb` · `sha256:9c713be4fc872b46e66d812e033708b51b4e385c81aa45f7b1798459a677343d`
- Locator: README News, lines 176-188
- License: Apache-2.0 · metadata and quote-minimal README spans

- Source: [DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/html/2506.11763v1)
- Work: `em:dossier-source-work:sha256:2a2c79463c9697b114728426edb183b5881f8b742ea6febfc272f73db58c1a14`
- Edition: `em:dossier-edition:sha256:c0f1d83abed0c9adcc86259b5bfe38e3f18b290f25227b5b9e77c99322fe0b2c` · `sha256:023aac7d22e7ca1230e2ffa237a4ea24c9f69f8b27653ddb6a494d4a30e4331c`
- Span: `em:dossier-span:sha256:17193f2aab7106eece041bcabda963d9bd172d798792efc058b6a93c9bd4cb0c` · `sha256:43c892b582d42a335c982009f97347b5bf2ae2799cc3e478c53504f4d01e4388`
- Locator: p.17, Appendix C, lines 704-708
- License: CC BY 4.0; unknown · quote-minimal attributed spans

### Support

DeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:6cf2883b02dd404844a662bae3367ab60d91f6847160a3df9c6ec6b287dc0858` · `sha256:bb4f28cd0f4527c633bcfacf51b18ced7387c4a2f266c59d41de7ee899b24b85`
- Locator: page 14, Table 3
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b` · `sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d`
- Span: `em:dossier-span:sha256:e748230bbc2f6b31433eaba4536917834c454cb73a7f18e8d9b975e2cac447b5` · `sha256:a0eeca5c86f9c7aeda9167f626a8cb0ed56dc87ea1ac5ca8c4a7a0add76e7625`
- Locator: Abstract, line 46
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:b9cbe6b1d35e74f3c8133d2be41a47d9361a26958e9f644dafd97e9e5acb516a` · `sha256:2ce91759797b86455a478a9ced58ec59d032e1a32a1c964347f96be4ab9e2054`
- Locator: pp.4-7, Sections 3.1.1 and 3.2
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:c7779c770aa24aa5925b5fe6e66f968c733e8522d30bcdc4290bc9409bbe8b5b` · `sha256:07a6f7e0282d888bc9157436b17fbde6286fb934b0efafd23399eebb7a22fa5d`
- Span: `em:dossier-span:sha256:8c1ac40e63a414c1a41cd0842632677269c8800967f22ef8247b6fb36f6242ea` · `sha256:5cc83ab699cb1342475b407e34d33fc106d64b5d9270e5bb9445d80586ebabc2`
- Locator: Section 3.2, lines 163-171
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:06ea026aae360f700369f4b1f658c4ca85e8acce431923cd7978ff2358b1ec05` · `sha256:af76864a50c641f8d693216d3b7ebe2a414eebf75467305c1e1a6946b6746590`
- Locator: Official ICLR abstract
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:12ab62355d178256bbaed6cdf5cbf6ae150410d9364f98cced347b54161fa5e8` · `sha256:a38c9ac275570f2b78dd1e1bf7663579e35f9037131e12d721dfb4c6f005683d`
- Locator: page 4, source scraping
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:2aa386b3653a99db3b670f783272056371480ed287dc1f62a51d3d587d680996` · `sha256:75a9212dc448c15ffc1e4a36bdf190728977470661bd9f45be534010704faa25`
- Locator: p.14, Table 3
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:a575506a3cca26acda376aa922fdb0e269767f71b01f502a7c50f3d4492e6f03` · `sha256:d2f6a4561a938ee913dadde2fc4a80c32c492fbcb3b47f236c9639375843ae8d`
- Locator: Paper section 3.1.1, factual-support validation
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:533550e262b8691ce9ed33aa0e89a7bf13d89c04c6002e1855ed1e767f556e3a` · `sha256:cf32a686cec566152749e3a4c643fa29aa115099f9eba77c0ca113c48264f769`
- Locator: printed page 10; PDF page index 10/22; Section 4 Results; Deep Research Agents paragraph
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

- Source: [DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence](https://arxiv.org/html/2509.04499v1)
- Work: `em:dossier-source-work:sha256:f58addec09e622c4c6457ae395e065ad2c0e9bc92035131100c8cb15e3a9d6d1`
- Edition: `em:dossier-edition:sha256:1a587168a59e161ea083b7b289339c5f41444a26f2f82a2eb8acb43b27367f70` · `sha256:0e0c4dd57c80c2ab32707ec1b17eda3ea21fc75ec4ce5191875446912649b545`
- Span: `em:dossier-span:sha256:cf4e063e6f297fe901b342752b67850d84cc3e8e21ace2ef8d5ea4f3fd93091c` · `sha256:e020a09063eee7aa1eb70663cf3e8bd7a8162952f5ae65d959d57ef74ab855f6`
- Locator: printed page 9; PDF page index 9/22; Table 1; Gemini (DR) column; %Citation Accuracy row
- License: arXiv non-exclusive distribution license; unknown · metadata, digest, and quote-minimal attributed spans only

### Support

Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:d4e2f35392b064365ad867f33ead8636bb8bbf3af03c0141585e5ae56c868090` · `sha256:a4386d56325583893d806fdd15ccd475d5096f0b445b1b9f52aeba253c910987`
- Locator: Section 3.3.1-3.3.3
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:ea33bc17365ed01ed68ab06f33c4fd2020dc8e6947659e4e9629b44af4aef635` · `sha256:be05510d58cc0231bf968f29c70b3d669a82d5cd0c74a778a5f8c4e035bf45b4`
- Locator: Abstract, lines 47-48
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:49ee66722d6d523dd485f05c796181af55736c3ceab2a235c5c3b40cca95e439` · `sha256:44d4c6941f2920e907de1e4cd22aeb9a725f24cc19432723754e316bd6f5cc49`
- Locator: Abstract and Table 1
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:8cf2ff4c32b2dac6d6dfc7684dd2bf7dca9d40fe7b7b36d50dfdcddad067bff3` · `sha256:e35ca66d12349d7290a6dbe7449bf6da173eb0671765644b581713528b316c74`
- Locator: Abstract and Table 1
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:237508b682de825711f61762deb5ce7661d5d197e9fde3aed5f5e80f05fdcdb5` · `sha256:d34096a9d9b1fbb43ef50ba503c3894607c008b563e1e1a5d1643cbc23c5d447`
- Locator: Abstract, lines 45-49
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents](https://arxiv.org/html/2605.06635v1)
- Work: `em:dossier-source-work:sha256:19d52572cf9e426b68cba35f9dd627d014ff71a978756c5b50103c2aabe554ce`
- Edition: `em:dossier-edition:sha256:601f0e4510819607265b94c8261884c072d559e63de9f71dbfc6bca5d33c5e29` · `sha256:90da1e5adccef8448150bd162daa5cdc08d01bfab5142d92de6aee947a79e58b`
- Span: `em:dossier-span:sha256:2bd67a5caa606894f4301ca1185c88a4dde02d3c5c4a2cf311463a0f39332443` · `sha256:da34a8e2a1d4e190d3b2095d79886b3cbf7edb2c63874189aa381302d2c2924e`
- Locator: Section 4.4, line 217
- License: CC BY 4.0 · quote-minimal attributed spans

### Support

The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.

- Source: [Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents](https://arxiv.org/html/2604.03173v1)
- Work: `em:dossier-source-work:sha256:ebddaffccfcd23869d5dd243d9f4272b520252f96cd784ed4058bacef7f764aa`
- Edition: `em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f` · `sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7`
- Span: `em:dossier-span:sha256:179697ec62048a6465fa50dda34ff54530e700430013a878201b9dadde405cc0` · `sha256:ee1de7374c67e208aa5d1331c20a08c8c83418309f1bd768d380620bfdc27901`
- Locator: pp. 1-2, lines 64-95
- License: CC0 1.0 · quote-minimal attributed spans

- Source: [Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents](https://arxiv.org/html/2604.03173v1)
- Work: `em:dossier-source-work:sha256:ebddaffccfcd23869d5dd243d9f4272b520252f96cd784ed4058bacef7f764aa`
- Edition: `em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f` · `sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7`
- Span: `em:dossier-span:sha256:73d72803976f4bb5529880a7185545e8b0ce5581409b9f64ba6ace3ac07c4e2c` · `sha256:3a14d6f574bed17e76ef8b560bbca0aa27b3a50918e49461433e34ec7770a4c9`
- Locator: page 3, Section 3.3
- License: CC0 1.0 · quote-minimal attributed spans

- Source: [Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents](https://arxiv.org/html/2604.03173v1)
- Work: `em:dossier-source-work:sha256:ebddaffccfcd23869d5dd243d9f4272b520252f96cd784ed4058bacef7f764aa`
- Edition: `em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f` · `sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7`
- Span: `em:dossier-span:sha256:57bdc827daee408c0042ec10327966b2b898d525449bee7ecbce1e02330f7f5f` · `sha256:b240bd208cb4c322f3c9c242d01e13da53b3fd2a64e41b8b0f26653032e0ee54`
- Locator: Section 3.3 'URL extraction and classification'
- License: CC0 1.0 · quote-minimal attributed spans

- Source: [Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents](https://arxiv.org/html/2604.03173v1)
- Work: `em:dossier-source-work:sha256:ebddaffccfcd23869d5dd243d9f4272b520252f96cd784ed4058bacef7f764aa`
- Edition: `em:dossier-edition:sha256:89224c483a2c88c51351393bd2ae482b8f2642a27b621c5a7a1401b05fe86b2f` · `sha256:37bffd052d581c627c32d65958b6037bebdc9a9c4cf1bef8376499fb0caa5bb7`
- Span: `em:dossier-span:sha256:1adfd92244c349596074ef5bf78e7edf7297e8966c13f429410a0f5871d89e45` · `sha256:9741316a5669b3b6e1e3022dfe4bfc23f53569a2839c46889423478a0b3c058f`
- Locator: Section 3 Experimental setup; Section 3.1 Datasets; section#S3.SS1; paragraph p#S3.SS1.p1.1
- License: CC0 1.0 · quote-minimal attributed spans

### Support

ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:bf94ce1ab912090b1db248bca0eae62fb77c0f4f12f2fffca9cafaf6524deadd` · `sha256:4947edaf643bd0a234e0ca914d920a1a5b068d18b017dcfe6d6103c466d422ca`
- Locator: Limitations, lines 259-261
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:e896138fd05fe0904311d35606553601c77784112c448651ce822a85e31ed99c` · `sha256:ac4714fdd34ead02999923441a133a798a67254f0449f293905ef5b1d1caffab`
- Locator: Section 4, line 186
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:d87df9d6e418184717863f7c621ddc6c16838dc569a1ab2a558421fbc20eeaa9` · `sha256:4545fdd46cab2ffc151c9ddbfaab08b82862b68eebf5205cda180b5875a3aa81`
- Locator: Section 3.3, line 155
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:3c149d37c88677f94f1715b536ad1808c4e476392f7e02212adbc723e30d8e27` · `sha256:e45fbf61feb38fc8de32433d439fdb8544133e905103b884c92327930a929b0a`
- Locator: p. 5, lines 315-329
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:5da64a3b244d7aa2f4baebee7a32be87f2f4255de92db8287d1b11a2a396932d` · `sha256:153b2f37a8662ac0dbbd7fe51bcbb291afd5d227c4da3f4c6463274c63a11fb2`
- Locator: Section 2.2 'Cited statements'
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:f3bc53b3bd401949fea37755e5943e417ceb79e044727a73f7219cfc824ab859` · `sha256:397d217cf5022ad78aa1072c2e82bcf3197e5a87cf4ef96ea81a78e1ad39b384`
- Locator: Settings and metrics, lines 132-139
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:abc076cdbec4977f668ce36fa6e5edcfebfd81c18c90a4f7913fbb0f8d5fe33a` · `sha256:b55cc3d05dbba5d160b75515ce407128fe2e2cf8c2d679517e79309d6ed6ab40`
- Locator: Section 3.1 'Setttings'
- License: CC BY 4.0 · quote-minimal attributed spans

- Source: [ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks](https://arxiv.org/html/2508.15804v1)
- Work: `em:dossier-source-work:sha256:a3c37defe96fa413eeb6391ba84067bec7456d35921aaf1b38c7ad003078fb54`
- Edition: `em:dossier-edition:sha256:c674e09fbcbadea10ca1f9c3bb17f5573f020fbc209410c85f3c8f56d53255b6` · `sha256:e13193419b2e087faabf99c51b27acf1a58919072019e4197c41c279354eb2f8`
- Span: `em:dossier-span:sha256:5ea28918cbd966b0fc14072426f2b7132bfe47f03d327e0b6eaee0629c3740ca` · `sha256:9b1bf034e823976ac267434e9c67ee7e4038da04f3d2582a0de7d6d729fc6916`
- Locator: Table 1, OpenAI Deep Research row and cited-statements Match Rate column
- License: CC BY 4.0 · quote-minimal attributed spans

## Complete count ledgers

### Captured reports (8)

- `V2-SOL-01` — V2-SOL-01 — completed
- `V2-SOL-02` — V2-SOL-02 — completed
- `V2-SOL-03` — V2-SOL-03 — completed
- `V2-SOL-04` — V2-SOL-04 — completed
- `V2-TERRA-01` — V2-TERRA-01 — completed
- `V2-TERRA-02` — V2-TERRA-02 — completed
- `V2-TERRA-03` — V2-TERRA-03 — completed
- `V2-TERRA-04` — V2-TERRA-04 — completed

### Citation occurrences (48)

- `V2-SOL-01:s1_keplinger` — V2-SOL-01:s1_keplinger — unresolved
- `V2-SOL-01:s1a_keplinger_data` — V2-SOL-01:s1a_keplinger_data — unresolved
- `V2-SOL-01:s2_drbench` — V2-SOL-01:s2_drbench — unresolved
- `V2-SOL-01:s2a_drbench_repo` — V2-SOL-01:s2a_drbench_repo — matched-exact-span
- `V2-SOL-01:s3_deeptrace` — V2-SOL-01:s3_deeptrace — unresolved
- `V2-SOL-01:s4_cited_not_verified` — V2-SOL-01:s4_cited_not_verified — unresolved
- `V2-SOL-01:s5_url_health` — V2-SOL-01:s5_url_health — unresolved
- `V2-SOL-02:S1` — V2-SOL-02:S1 — unresolved
- `V2-SOL-02:S2` — V2-SOL-02:S2 — matched-exact-span
- `V2-SOL-02:S3` — V2-SOL-02:S3 — unresolved
- `V2-SOL-02:S4` — V2-SOL-02:S4 — unresolved
- `V2-SOL-02:S5` — V2-SOL-02:S5 — unresolved
- `V2-SOL-02:S6` — V2-SOL-02:S6 — matched-exact-span
- `V2-SOL-02:S7` — V2-SOL-02:S7 — unresolved
- `V2-SOL-03:s1` — V2-SOL-03:s1 — unresolved
- `V2-SOL-03:s2` — V2-SOL-03:s2 — matched-exact-span
- `V2-SOL-03:s3` — V2-SOL-03:s3 — unresolved
- `V2-SOL-03:s4` — V2-SOL-03:s4 — unresolved
- `V2-SOL-03:s5` — V2-SOL-03:s5 — matched-exact-span
- `V2-SOL-03:s6` — V2-SOL-03:s6 — matched-exact-span
- `V2-SOL-03:s7` — V2-SOL-03:s7 — matched-exact-span
- `V2-SOL-04:S1` — V2-SOL-04:S1 — matched-exact-span
- `V2-SOL-04:S2` — V2-SOL-04:S2 — matched-exact-span
- `V2-SOL-04:S3` — V2-SOL-04:S3 — matched-exact-span
- `V2-SOL-04:S4` — V2-SOL-04:S4 — matched-exact-span
- `V2-SOL-04:S5` — V2-SOL-04:S5 — matched-exact-span
- `V2-SOL-04:S6` — V2-SOL-04:S6 — unresolved
- `V2-SOL-04:S7` — V2-SOL-04:S7 — matched-exact-span
- `V2-TERRA-01:s1_deepresearchbench` — V2-TERRA-01:s1_deepresearchbench — unresolved
- `V2-TERRA-01:s2_reportbench` — V2-TERRA-01:s2_reportbench — unresolved
- `V2-TERRA-01:s3_researcherbench` — V2-TERRA-01:s3_researcherbench — unresolved
- `V2-TERRA-01:s4_cited_not_verified` — V2-TERRA-01:s4_cited_not_verified — unresolved
- `V2-TERRA-01:s5_urlhealth` — V2-TERRA-01:s5_urlhealth — unresolved
- `V2-TERRA-02:s1_deepresearchbench` — V2-TERRA-02:s1_deepresearchbench — unresolved
- `V2-TERRA-02:s2_liveresearchbench` — V2-TERRA-02:s2_liveresearchbench — unresolved
- `V2-TERRA-02:s3_deeptrace` — V2-TERRA-02:s3_deeptrace — unresolved
- `V2-TERRA-02:s4_researcherbench` — V2-TERRA-02:s4_researcherbench — unresolved
- `V2-TERRA-03:S1` — V2-TERRA-03:S1 — unresolved
- `V2-TERRA-03:S2` — V2-TERRA-03:S2 — unresolved
- `V2-TERRA-03:S3` — V2-TERRA-03:S3 — unresolved
- `V2-TERRA-03:S4` — V2-TERRA-03:S4 — unresolved
- `V2-TERRA-03:S5` — V2-TERRA-03:S5 — matched-exact-span
- `V2-TERRA-03:S6` — V2-TERRA-03:S6 — unresolved
- `V2-TERRA-04:S1_deepresearchbench` — V2-TERRA-04:S1_deepresearchbench — unresolved
- `V2-TERRA-04:S2_researcherbench` — V2-TERRA-04:S2_researcherbench — unresolved
- `V2-TERRA-04:S3_reportbench` — V2-TERRA-04:S3_reportbench — unresolved
- `V2-TERRA-04:S4_urlhealth` — V2-TERRA-04:S4_urlhealth — unresolved
- `V2-TERRA-04:S5_cited_not_verified` — V2-TERRA-04:S5_cited_not_verified — unresolved

### Distinct URL strings (30)

- `https://arxiv.org/abs/2604.03173` — https://arxiv.org/abs/2604.03173 — retrieved
- `https://arxiv.org/abs/2605.06635` — https://arxiv.org/abs/2605.06635 — retrieved
- `https://arxiv.org/abs/2607.08700` — https://arxiv.org/abs/2607.08700 — retrieved
- `https://arxiv.org/html/2506.11763` — https://arxiv.org/html/2506.11763 — retrieved
- `https://arxiv.org/html/2506.11763v1` — https://arxiv.org/html/2506.11763v1 — retrieved
- `https://arxiv.org/html/2507.16280` — https://arxiv.org/html/2507.16280 — retrieved
- `https://arxiv.org/html/2508.15804` — https://arxiv.org/html/2508.15804 — retrieved
- `https://arxiv.org/html/2508.15804v1` — https://arxiv.org/html/2508.15804v1 — retrieved
- `https://arxiv.org/html/2509.04499` — https://arxiv.org/html/2509.04499 — retrieved
- `https://arxiv.org/html/2509.04499v1` — https://arxiv.org/html/2509.04499v1 — retrieved
- `https://arxiv.org/html/2604.03173` — https://arxiv.org/html/2604.03173 — retrieved
- `https://arxiv.org/html/2604.03173v1` — https://arxiv.org/html/2604.03173v1 — retrieved
- `https://arxiv.org/html/2605.06635` — https://arxiv.org/html/2605.06635 — retrieved
- `https://arxiv.org/html/2605.06635v1` — https://arxiv.org/html/2605.06635v1 — retrieved
- `https://arxiv.org/pdf/2507.16280` — https://arxiv.org/pdf/2507.16280 — retrieved
- `https://arxiv.org/pdf/2508.15804` — https://arxiv.org/pdf/2508.15804 — retrieved
- `https://arxiv.org/pdf/2510.14240v5` — https://arxiv.org/pdf/2510.14240v5 — retrieved
- `https://arxiv.org/pdf/2604.03173` — https://arxiv.org/pdf/2604.03173 — retrieved
- `https://arxiv.org/pdf/2605.06635` — https://arxiv.org/pdf/2605.06635 — retrieved
- `https://data.mendeley.com/datasets/3s73z9zf3c/1` — https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
- `https://data.mendeley.com/datasets/3s73z9zf3c/2` — https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
- `https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf` — https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
- `https://github.com/Ayanami0730/deep_research_bench` — https://github.com/Ayanami0730/deep_research_bench — retrieved
- `https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035` — https://onlinelibrary.wiley.com/doi/10.1111/jdv.70035 — inaccessible
- `https://openreview.net/pdf?id=hQ0K2Hhq7H` — https://openreview.net/pdf?id=hQ0K2Hhq7H — inaccessible
- `https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/` — https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf` — https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf` — https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html` — https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved
- `https://pubmed.ncbi.nlm.nih.gov/40904191/` — https://pubmed.ncbi.nlm.nih.gov/40904191/ — inaccessible

### Resolving URL roots (27)

- `https://arxiv.org/abs/2604.03173` — https://arxiv.org/abs/2604.03173 — retrieved
- `https://arxiv.org/abs/2605.06635` — https://arxiv.org/abs/2605.06635 — retrieved
- `https://arxiv.org/abs/2607.08700` — https://arxiv.org/abs/2607.08700 — retrieved
- `https://arxiv.org/html/2506.11763` — https://arxiv.org/html/2506.11763 — retrieved
- `https://arxiv.org/html/2506.11763v1` — https://arxiv.org/html/2506.11763v1 — retrieved
- `https://arxiv.org/html/2507.16280` — https://arxiv.org/html/2507.16280 — retrieved
- `https://arxiv.org/html/2508.15804` — https://arxiv.org/html/2508.15804 — retrieved
- `https://arxiv.org/html/2508.15804v1` — https://arxiv.org/html/2508.15804v1 — retrieved
- `https://arxiv.org/html/2509.04499` — https://arxiv.org/html/2509.04499 — retrieved
- `https://arxiv.org/html/2509.04499v1` — https://arxiv.org/html/2509.04499v1 — retrieved
- `https://arxiv.org/html/2604.03173` — https://arxiv.org/html/2604.03173 — retrieved
- `https://arxiv.org/html/2604.03173v1` — https://arxiv.org/html/2604.03173v1 — retrieved
- `https://arxiv.org/html/2605.06635` — https://arxiv.org/html/2605.06635 — retrieved
- `https://arxiv.org/html/2605.06635v1` — https://arxiv.org/html/2605.06635v1 — retrieved
- `https://arxiv.org/pdf/2507.16280` — https://arxiv.org/pdf/2507.16280 — retrieved
- `https://arxiv.org/pdf/2508.15804` — https://arxiv.org/pdf/2508.15804 — retrieved
- `https://arxiv.org/pdf/2510.14240v5` — https://arxiv.org/pdf/2510.14240v5 — retrieved
- `https://arxiv.org/pdf/2604.03173` — https://arxiv.org/pdf/2604.03173 — retrieved
- `https://arxiv.org/pdf/2605.06635` — https://arxiv.org/pdf/2605.06635 — retrieved
- `https://data.mendeley.com/datasets/3s73z9zf3c/1` — https://data.mendeley.com/datasets/3s73z9zf3c/1 — retrieved
- `https://data.mendeley.com/datasets/3s73z9zf3c/2` — https://data.mendeley.com/datasets/3s73z9zf3c/2 — retrieved
- `https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf` — https://deepresearch-bench.github.io/static/papers/deepresearch-bench.pdf — retrieved
- `https://github.com/Ayanami0730/deep_research_bench` — https://github.com/Ayanami0730/deep_research_bench — retrieved
- `https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/` — https://pmc.ncbi.nlm.nih.gov/articles/PMC13109748/ — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf` — https://proceedings.iclr.cc/paper_files/paper/2026/file/465f22be10e07b301c6ed58f0472f704-Paper-Conference.pdf — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf` — https://proceedings.iclr.cc/paper_files/paper/2026/file/ad08767706825033b99122332293033d-Paper-Conference.pdf — retrieved
- `https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html` — https://proceedings.iclr.cc/paper_files/paper/2026/hash/ad08767706825033b99122332293033d-Abstract-Conference.html — retrieved

### Source works (11)

- `work-citation-verifier-benchmark-2d5e94336b` — Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution — examined-source-work
- `work-cited-not-verified-3825b25622` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — examined-source-work
- `work-deepresearch-bench-paper-8ef36010d2` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — examined-source-work
- `work-deepresearch-bench-repository-1e475c7631` — Ayanami0730/deep_research_bench — examined-source-work
- `work-deeptrace-650d66cfcb` — DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — examined-source-work
- `work-keplinger-dermatology-audit-639f9ea37a` — Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — examined-source-work
- `work-keplinger-supplement-1a03034fb3` — Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — examined-source-work
- `work-liveresearchbench-61e7da557b` — LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — examined-source-work
- `work-reportbench-dca823c910` — ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — examined-source-work
- `work-researcherbench-82729ab60f` — ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — examined-source-work
- `work-url-health-f64342d493` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — examined-source-work

### Examined editions (14)

- `edition-citation-verifier-arxiv-v1-acfd4abab7` — Quote-minimal projection of edition:citation-verifier-arxiv-v1 — examined-edition
- `edition-cnv-arxiv-v1-f24001c50a` — Quote-minimal projection of edition:cnv-arxiv-v1 — examined-edition
- `edition-deeptrace-arxiv-v1-70a6e720c1` — Quote-minimal projection of edition:deeptrace-arxiv-v1 — examined-edition
- `edition-deeptrace-iclr-2026-b2b153eaa6` — Quote-minimal projection of edition:deeptrace-iclr-2026 — examined-edition
- `edition-drbench-arxiv-v1-7afa5ade36` — Quote-minimal projection of edition:drbench-arxiv-v1 — examined-edition
- `edition-drbench-iclr-2026-d00a3dcdb6` — Quote-minimal projection of edition:drbench-iclr-2026 — examined-edition
- `edition-drbench-repo-main-469cce5-4f8964f3e3` — Quote-minimal projection of edition:drbench-repo-main-469cce5 — examined-edition
- `edition-keplinger-supplement-v1-f6ddb4c04a` — Quote-minimal projection of edition:keplinger-supplement-v1 — examined-edition
- `edition-keplinger-supplement-v2-01e44c09d8` — Quote-minimal projection of edition:keplinger-supplement-v2 — examined-edition
- `edition-keplinger-vor-2025-10ba10f860` — Quote-minimal projection of edition:keplinger-vor-2025 — examined-edition
- `edition-liveresearchbench-arxiv-v5-f5644763e5` — Quote-minimal projection of edition:liveresearchbench-arxiv-v5 — examined-edition
- `edition-reportbench-arxiv-v1-25888dc44b` — Quote-minimal projection of edition:reportbench-arxiv-v1 — examined-edition
- `edition-researcherbench-arxiv-v1-39a8b256d7` — Quote-minimal projection of edition:researcherbench-arxiv-v1 — examined-edition
- `edition-url-health-arxiv-v1-85fb934b53` — Quote-minimal projection of edition:url-health-arxiv-v1 — examined-edition

### Accepted exact spans (72)

- `span-015bb68c904bf225e84c82b1acffb3cc38a4a087d7b0f79d-4c454092b1` — Evaluation methodology, lines 144-149 — matched-exact-span
- `span-084f11e57cf779850ea87aaf30ba1277438ff9efcb533fd2-fe72b27cfb` — Appendix A.1 'Limitations' — matched-exact-span
- `span-092745c7db6407d4d52927642741cde229fad05a5347bb18-7b04227f8f` — Main text, paragraph beginning 'For their high performances in generating reference lists' — matched-exact-span
- `span-0abbab3d6269d58eb76d8e90eff1e4c27512cf46d9b8e5ff-54e5759a8b` — repository README, Overview — matched-exact-span
- `span-0d4f357f2a5565bdaea01a1e537650958d71216868d1a2b0-9b0af35d76` — Table 1, lines 177-182 — matched-exact-span
- `span-13cb0ddb087b82988b7445a9cfb02562f9c044ef4f244e8b-87c3b42c68` — Section 3.2, lines 144-146 — matched-exact-span
- `span-196c54d6212fbc0c054dbf8e8467f2388c1f788bb93f4b1a-d460e927f6` — Table 1, Perplexity Deep Research citation-accuracy cell — matched-exact-span
- `span-1db12b68cfeb4c5f5a96fa9cd2e2bd46e8dfb4f096f73154-cea47f0c65` — Limitations, lines 259-261 — matched-exact-span
- `span-279db725380c133044944dccb379b75a396242219aa30437-569fc9e2fd` — page 16, FACT judge validation — matched-exact-span
- `span-2f311bfb451c6208101cb7583af37c338914fcdb10f5c75b-90e7764ce2` — Table 1, OpenAI Deep Research row; FACT columns — matched-exact-span
- `span-3015a109e7898cb720e426e9a57c526277b648c75f1fb4b1-745d84ff44` — Section 4, line 186 — matched-exact-span
- `span-35927851edd2e7e5e0f60d98498940f88304ba99bf1f85a0-2fe586436b` — page 14, Table 3 — matched-exact-span
- `span-39a9931f05ea356fe904ed15768fa08c66440b1fcb20f98e-191c24bbb2` — Table 1, ChatGPT all-correct subtotal — matched-exact-span
- `span-417182b4c15ac550469e435a02aab1e91b395f72f186312d-e9e2c4a70c` — Section 4.3, line 212 — matched-exact-span
- `span-46c392bc942cd88d525d74298e6bc6ebd9aeaec533ddfb7b-774f5b996e` — main text, paragraph beginning 'For their high performances'; Figure 1 — matched-exact-span
- `span-55de6a6ca9414a252592e9f77544a9259533dd713ce1f12f-0826216ffc` — pp. 1-2, lines 64-95 — matched-exact-span
- `span-59d0da64f738e1d8745f2bff3663eb89e273f93fe501afe3-ba69fc58eb` — Section 2 definition discussion and Section 3.3 — matched-exact-span
- `span-5cd9ba4ca72da10cf51c0beddc6ea2765c74721bdd526a67-ed8f789629` — Limitations, lines 219-224 — matched-exact-span
- `span-6885cf2524462022475b5829da081b2d35fcfe7463ca6ad0-3b1a122faf` — Section 1 contributions — matched-exact-span
- `span-6c1d327e58ef001533ea1c11010e974436bbeafbde138efc-f6efd8d376` — Section 3.3, line 155 — matched-exact-span
- `span-6e961498ed0f81a950f775956587f741e0eb9f11640fd3cb-afa5b59854` — Section 4.3, Tables 2–3 — matched-exact-span
- `span-6f79c3c824ccf889d73e234335661e5fca68d90f8c78ee97-dbf7848a72` — dataset page, Description — matched-exact-span
- `span-732c03e69c6e142a492275cee5d1e2d474c75422c9b84aaa-f7675998cc` — Section 3.3.1-3.3.3 — matched-exact-span
- `span-74f6acafbb4c54d556a572460bf1b45bacd1e293f4aa16cc-5ee15c82c4` — Abstract, lines 47-48 — matched-exact-span
- `span-76e9bd59b31c7ebd2ed7a2186e59753a79d3c730e0d7b4fb-9c98d99242` — Appendix D, Table 5 — matched-exact-span
- `span-794bdf7e0595799afbb0d55f3dc522077aed9615c59f1e35-43f2ba944d` — Abstract, line 16 — matched-exact-span
- `span-7a6046a756d6ec0315d330cf0817d062ea48498bdc9c90b3-c3a1e80029` — p. 5, lines 315-329 — matched-exact-span
- `span-7d7331942f9d0520e18e96206e7718e5187788be4d48cef9-dfdbc1d019` — page 4, Section 3.2, Support Judgment — matched-exact-span
- `span-886b50d94e3facabc487342cd2af5157cf956622ecfc3686-e7bdad0935` — Section 4.2.1, Citation Support Verification and Score Computation — matched-exact-span
- `span-8c8d9a6dc003bf0298542b85df3f6f08db631e68e89f606f-92731379f2` — Table 1, ChatGPT subtotal — matched-exact-span
- `span-8cce3d9ee1e1674ba1b93321dd9b99a55d834ca980451f8a-d2fceafe11` — Abstract, line 46 — matched-exact-span
- `span-912c931063109f84ea88aea34891dd5e0f9147d2176af258-4d566ca076` — Abstract and Table 1 — matched-exact-span
- `span-92eb7cf19745e24190dda84c77d590941134d09f06805c0f-88afe43377` — README FACT, lines 240-248 — matched-exact-span
- `span-94ad583514aa7e04b8d5f7b8f9cdda6fdffe8ca6712f033d-63c1ba24eb` — Dataset description — matched-exact-span
- `span-9c7a4c186cc85f185aa293a7c5a46c08f9dbbc1a9e63c449-feff07a6e2` — Section 2.2 'Cited statements' — matched-exact-span
- `span-9cb1700c08368feb234207045354cbf559d6ed80f2b0bae8-f898722eb5` — pp.4-7, Sections 3.1.1 and 3.2 — matched-exact-span
- `span-9dd6eb46529e77a62748093528e52685fbd92d46a756aef7-96d75d9e2b` — Section 3.2, lines 163-171 — matched-exact-span
- `span-9fe42cb48702d8d96ebbbc0309692214ece6ec495c447c06-80da541506` — Section 3.2, paragraphs 'Statement-URL Pair Extraction and Deduplication' and 'Support Judgment' — matched-exact-span
- `span-a08468cb62f2feb21d9cbb95175ce00f54592bd83573cc91-03f65cfa45` — Official ICLR abstract — matched-exact-span
- `span-a15ab02e9fcb23df02fb03d28967d4220ddf65ad380d2230-ddf85fa359` — page 4, source scraping — matched-exact-span
- `span-a18e778fac3d49e81167f05e09fbc361e91af8c0c01b0fec-ca1b18e206` — page 3, Section 3.3 — matched-exact-span
- `span-aa1534ce2e177153ab85e512deb8d064d14ec1eae16d8032-2299096a09` — p.14, Table 3 — matched-exact-span
- `span-b0a39c963cdfdc5f9852cfc9f7e33e05d31232bd1d899885-adf9545222` — p. 4, lines 150-159 — matched-exact-span
- `span-b38a2af5982b4c85bc205d2f533a23ed3a8f40a49b651a5f-27395a6291` — Main text, paragraph beginning 'To address the gap' (search-indexed primary full text) — matched-exact-span
- `span-b68153666e4e12aee8dc887fef0a61fa5f44b5f680a0b7eb-769c2cc17d` — Section 5.1 Results — matched-exact-span
- `span-b6a6ba48a2c4e352c8565b671a15a90400739f2aeefda8a6-e8232daf27` — Main text, following paragraph (search-indexed primary full text) — matched-exact-span
- `span-b83996e3b5fdf878e04d6d41d0e7a1eebee9fdd9bbccf2ae-820a3eab86` — Appendix C, line 365 — matched-exact-span
- `span-b903dc2284e1b4d80bb6ef948d79bbdbc7d7447e871281b8-6ef6234316` — body, paragraph immediately after Table 1 — matched-exact-span
- `span-ba6b8a8b7d9121ebb05d594176ceed3d1e9dbc637876783a-d85147b6b0` — Settings and metrics, lines 132-139 — matched-exact-span
- `span-bb16d5f06abe8641d9ea6694e549aa15e240dd88946aa47a-5946b6e6b7` — Abstract and Table 1 — matched-exact-span
- `span-bb62928b5e0cb4e373b2f6bfeb579d8564d05da5588f8b15-771310f4ef` — Section 3.3 'URL extraction and classification' — matched-exact-span
- `span-c27aaca02bda06ed764be53351158fc862af9d4a556d3e82-104e7cda52` — page 5, Section 4.1 — matched-exact-span
- `span-c3dfc98566c27561eb7267972e00e3e1ba6a399a3ff5d002-37f40d3260` — page 5, Section 3.3.3 — matched-exact-span
- `span-c7dbd1d9e09980703ecbee8161594e4b426c1af63117b286-153d26f046` — Section 5 'Limitations' — matched-exact-span
- `span-cd21c450c0d33df20a2b1a539d5725347db1b5bab7d1e986-2e3e0c3ac7` — Abstract, lines 45-49 — matched-exact-span
- `span-d638c8b21f0e4a0f840880b6ffe0217ce1623758cc09fdcc-8585f6d39b` — Section 4.3, Tables 2-3 and following paragraph — matched-exact-span
- `span-d880991ad784390029d49f84949d083c4789d67d45893d87-c6ef8a5c13` — p.33, Appendix E — matched-exact-span
- `span-ddf4320a6d302c13d28b6225eb7e79ac15583349e1b449d4-f89bb46179` — Section 3.1 'Setttings' — matched-exact-span
- `span-deada8bacb987a47f2a42e52fbfeb5c1beb3d250c0be79e2-0e91f5cacc` — Section 4.4, line 217 — matched-exact-span
- `span-e223640cb0ed12c87ca1d2406f3276f30a3b8d2017dd4dc1-3cfecd76bf` — Table 1, OpenAI Deep Research row and cited-statements Match Rate column — matched-exact-span
- `span-e2c9f6220843a04665f5d1cd142ac15f04211ace5c579fa6-c5e0476df0` — Paper section 3.1.1, factual-support validation — matched-exact-span
- `span-e4a9f30effc031388ed5c9993dc6f071903bec5e3f1998db-5d95269b76` — pp.17-18, Appendix E, equations 4-6 — matched-exact-span
- `span-e54ee1457854a9a5e46f7c7c2e66ac6a5a69b963f8ee3450-8f875cfe94` — Abstract, line 16 — matched-exact-span
- `span-e5bcf0241a1aeae7d9a92755b1bcf0adeeea93a8922f22f8-77a5ff4cc0` — README News, 2026-05-11 evaluator migration notice — matched-exact-span
- `span-e6f447129d9a662fb202d4ac22e16822bacfdbcbcdb7afe3-f7abe2194c` — p.23, Appendix C, Citation Accuracy — matched-exact-span
- `span-ea9cb2e1b0078c0857b789d61d0a0d26801cf66891864705-2f19633d83` — Section 5.1 and Appendix E, Table 4 — matched-exact-span
- `span-eaf37bef7b8dc35f3e00b85d3e58ad872dc0ea604a062c8c-fa09300f85` — p.4, Section 3.2, lines 148-155 — matched-exact-span
- `span-ef637b38f36195bea09205d532e8744daf3e257f93d97144-26007673a3` — README News, lines 176-188 — matched-exact-span
- `span-f05b3189dc3fafd45b5bde643a91fd0616832928215a37a2-47f72e77e4` — Paragraph after Table 1 — matched-exact-span
- `span-f35c9565f2fd047f6e7272dc608bbf6a63813a717bad5354-e6a07f05b6` — Appendix C, line 365 — matched-exact-span
- `span-f96ad5daa81153a3689f52f5777fedfd70e8a87d44e7a1c9-3da72b8748` — Section 4.2.1 'Citation Support Verification' and Section 4.2 'Score Computation', equations (2)-(3) — matched-exact-span
- `span-f98f1274859ba98cc538e3be6d11d345a5382cfcb9516aa1-d60b765c48` — p.17, Appendix C, lines 704-708 — matched-exact-span

### Candidate warrants (7)

- `relation-cnv-link-relevance-support-gap-165443d270` — Cited but Not Verified operationalized link access, topical relevance, and factual support separately and reported materially different rates.
- `relation-cnv-search-depth-ablation-9c01321092` — Within the Cited but Not Verified harness, the 2-to-150-call ablation reduced reported Fact Check scores for two setups without establishing a general causal law about search depth.
- `relation-deepresearch-bench-fact-a3f61d7a3a` — DeepResearch Bench evaluated binary support for deduplicated statement-URL pairs; numeric values are edition-specific.
- `relation-deeptrace-support-variation-1a84ef7c6f` — DeepTRACE measured statement-source support and found between-system variation; its Gemini value is internally inconsistent between table and prose.
- `relation-keplinger-metadata-versus-support-c23e105c8b` — In one five-system dermatology-review audit, reference identifiability and metadata correctness did not establish sentence-level claim-citation concordance.
- `relation-reportbench-match-rate-6964e41a69` — ReportBench compared cited statements with retrieved cited-page content and reported sub-100-percent semantic match rates in its bounded survey-task setting.
- `relation-url-health-resolution-0b26fb45c7` — The URL-health study measured HTTP resolution and Wayback presence, not semantic claim support, on outputs including reused DeepResearch Bench material.

### Pending warrants (4)

- `relation-citation-verifier-calibration-5b40082faa` — The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
- `relation-liveresearchbench-e1-e2-e3-5fb5d671f7` — The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
- `relation-researcherbench-faithfulness-groundedness-c84dcaee54` — The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.
- `relation-url-health-correction-loop-eb8a6384c6` — The captured spans do not semantically close the normalized proposition; the warrant remains pending and receives no credit.

### Unresolved citations (34)

- `V2-SOL-01:s1_keplinger` — Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
- `V2-SOL-01:s1a_keplinger_data` — Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
- `V2-SOL-01:s2_drbench` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-SOL-01:s3_deeptrace` — DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
- `V2-SOL-01:s4_cited_not_verified` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
- `V2-SOL-01:s5_url_health` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
- `V2-SOL-02:S1` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
- `V2-SOL-02:S3` — Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
- `V2-SOL-02:S4` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-SOL-02:S5` — ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
- `V2-SOL-02:S7` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
- `V2-SOL-03:s1` — Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
- `V2-SOL-03:s3` — DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
- `V2-SOL-03:s4` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
- `V2-SOL-04:S6` — Supplementary materials of the article: Assessment of Deep Research for Dermatology Literature Reviews: Deep Concern Over the Hype — unresolved
- `V2-TERRA-01:s1_deepresearchbench` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-TERRA-01:s2_reportbench` — ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
- `V2-TERRA-01:s3_researcherbench` — ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
- `V2-TERRA-01:s4_cited_not_verified` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
- `V2-TERRA-01:s5_urlhealth` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
- `V2-TERRA-02:s1_deepresearchbench` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-TERRA-02:s2_liveresearchbench` — LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild — unresolved
- `V2-TERRA-02:s3_deeptrace` — DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence — unresolved
- `V2-TERRA-02:s4_researcherbench` — ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
- `V2-TERRA-03:S1` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved
- `V2-TERRA-03:S2` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-TERRA-03:S3` — ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
- `V2-TERRA-03:S4` — Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
- `V2-TERRA-03:S6` — PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
- `V2-TERRA-04:S1_deepresearchbench` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-TERRA-04:S2_researcherbench` — ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry — unresolved
- `V2-TERRA-04:S3_reportbench` — ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks — unresolved
- `V2-TERRA-04:S4_urlhealth` — Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — unresolved
- `V2-TERRA-04:S5_cited_not_verified` — Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents — unresolved

### Unsupported or force-raised claims (20)

- `V2-SOL-01:r3_deepresearch_bench_fact` — V2-SOL-01:r3_deepresearch_bench_fact — no-credit
- `V2-SOL-01:r4_deeptrace_support` — V2-SOL-01:r4_deeptrace_support — no-credit
- `V2-SOL-02:R3_deepresearch_bench_fact` — V2-SOL-02:R3_deepresearch_bench_fact — no-credit
- `V2-SOL-02:R5_deeptrace_audit` — V2-SOL-02:R5_deeptrace_audit — no-credit
- `V2-SOL-03:r3_deepresearch_bench_fact` — V2-SOL-03:r3_deepresearch_bench_fact — no-credit
- `V2-SOL-03:r4_deeptrace_support` — V2-SOL-03:r4_deeptrace_support — no-credit
- `V2-SOL-03:r8_search_depth_ablation` — V2-SOL-03:r8_search_depth_ablation — no-credit
- `V2-SOL-03:r9_verifier_calibration` — V2-SOL-03:r9_verifier_calibration — no-credit
- `V2-SOL-04:R1` — V2-SOL-04:R1 — no-credit
- `V2-SOL-04:R2` — V2-SOL-04:R2 — no-credit
- `V2-SOL-04:R3` — V2-SOL-04:R3 — no-credit
- `V2-SOL-04:R4` — V2-SOL-04:R4 — no-credit
- `V2-SOL-04:R5` — V2-SOL-04:R5 — no-credit
- `V2-SOL-04:R8` — V2-SOL-04:R8 — no-credit
- `V2-TERRA-01:r1_deepresearchbench_fact` — V2-TERRA-01:r1_deepresearchbench_fact — no-credit
- `V2-TERRA-01:r4_cited_not_verified_source_attribution` — V2-TERRA-01:r4_cited_not_verified_source_attribution — no-credit
- `V2-TERRA-02:answer` — V2-TERRA-02:answer — no-credit
- `V2-TERRA-02:r1_deepresearchbench_fact` — V2-TERRA-02:r1_deepresearchbench_fact — no-credit
- `V2-TERRA-02:r4_deeptrace` — V2-TERRA-02:r4_deeptrace — no-credit
- `V2-TERRA-03:R3` — V2-TERRA-03:R3 — no-credit

### Independently rejected claims (9)

- `V2-SOL-02:R5_deeptrace_audit` — V2-SOL-02:R5_deeptrace_audit — independently-rejected
- `V2-SOL-03:r3_deepresearch_bench_fact` — V2-SOL-03:r3_deepresearch_bench_fact — independently-rejected
- `V2-SOL-03:r8_search_depth_ablation` — V2-SOL-03:r8_search_depth_ablation — independently-rejected
- `V2-SOL-03:r9_verifier_calibration` — V2-SOL-03:r9_verifier_calibration — independently-rejected
- `V2-SOL-04:R1` — V2-SOL-04:R1 — independently-rejected
- `V2-SOL-04:R2` — V2-SOL-04:R2 — independently-rejected
- `V2-SOL-04:R3` — V2-SOL-04:R3 — independently-rejected
- `V2-SOL-04:R4` — V2-SOL-04:R4 — independently-rejected
- `V2-SOL-04:R8` — V2-SOL-04:R8 — independently-rejected

### Inaccessible carriers (3)

- `V2-SOL-02:S4` — DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents — unresolved
- `V2-SOL-03:s1` — Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved
- `V2-TERRA-03:S6` — PubMed record and Figure 1 caption for Assessment of Deep Research for dermatology literature reviews: Deep concern over the hype — unresolved

## Reproducibility identity

- Dossier: `em:dossier:sha256:cbd7a14096a956f642f5c76046d3b49ed648fbe6bf24144c992404a01415af82`
- Review receipt: `dd7f8ad5f760137d91346c3bf38b2bbfffbc7e5c2e74a8a987b76d857e4f244e`
- Catalog: `em:catalog:sha256:092898e1fe3d355761ab4cec653576926a8f5d31621ec7ce23dd60e9d19563ef`
- Frontier: `em:frontier:sha256:7e4a173112ef26422acf3ed9434c8b6849c4e011797e20fed6c0a9ca58a1e4c3`
- Accepted commit: `af081caa99fc08d3fabb914ff68f2e672a83bd5b`
- Content digest: `048c12622d9daca7cd009a7483c58697aa0674f77b12c6eca3721954ec1f3743`
