Repository object · documentation
Case 002 research candidate: agent citation lineage
Status: v1 transport preflight failed and is retained; the v2 packet passed EM-0026 source review; an EM-0029 candidate dossier now exists and awaits exact-head independent review; research only.
- Media type
text/markdown- Object ID
em:documentation:sha256:a16ca25bcf6ad3470b4f122f22424ee40683cf22c0c53f3f5b55510a89f15f42- Content digest
dc444f08c8799ebba4287ac75f7d71ed8697f79a6da7fc72d07e4736aae8fadc
Source content
Case 002 research candidate: agent citation lineage
Status: v1 transport preflight failed and is retained; the v2 packet passed EM-0026 source review;
an EM-0029 candidate dossier now exists and awaits exact-head independent review; research only.
Case 002 asks whether apparent agreement among research agents survives inspection of the sources
and exact spans behind their citations. The pilot is designed to teach a distinct lesson from Case
001: a report, URL, citation, source work, source span, and warranted proposition are different
units.
Frozen target
What empirical evidence published or publicly posted by 2026-08-22 measures whether citations
produced by deep-research agents resolve and actually support the claims made from them?
The dated cutoff prevents a fast-moving product comparison from masquerading as a timeless result.
The question is about public empirical evaluations, not private vendor performance or all agents.
Collection design
The active v2 protocol assigns eight context-isolated runs: four on gpt-5.6-sol and four on
gpt-5.6-terra, each at requested high reasoning effort. Every run receives the same frozen prompt
with no inherited conversation turns. Unavailable provider system, sampling, or retrieval settings
remain unknown.
Before v2 was frozen, the v1 transport preflight detected that one invocation received a prompt
with one extra character. Both started invocations were stopped before final-answer capture. The
failed v1 matrix is not being completed or repaired in place: its exact prompt identities, terminal
statuses, and zero admitted answers remain a negative preflight record. V2 is a separately
identified matrix, not a replacement run inside v1.
Different runtime profiles are not different epistemic observers. The matrix gives the project a
bounded set of outputs to inspect; it does not turn eight reports into eight evidence roots.
What was inspected
For each citation, later work must resolve:
- requested and final URL, redirect chain, retrieval status, media type, bytes, and digest where
- source work, edition, exact span, and license treatment;
- the proposition asserted by the agent and the narrower proposition the span actually supports;
- modality, causality, scope, time, population, metric, comparison-class, and numerical-strength
- shared model, prompt, retrieval, URL, source, span, upstream citation, data, method, and
an artifact can be captured;
changes; and
derivation dependencies.
The deterministic research-only ledger currently reports:
8 agent reports → 30 cited URL strings → 27 resolving URL roots → 11 source works → 14
examined editions → 72 matched exact-span roots → 7 author-candidate warrant roots → 0
independently confirmed warrant roots
Every number is computed from captured records and relations. The raw matrix contains 48 citation
occurrences, 127 source-span occurrences, and 52 result claims. Twenty claim occurrences are
unsupported or force-raised after review, including nine whose linked quote fragments do not
support their complete proposition. Thirty-four citation occurrences remain unresolved and
receive no automatic credit. Three cited carriers were inaccessible during
fresh readback: two returned HTTP 403 and PubMed returned a cookie interstitial under HTTP 203. No
URL was malformed or found to identify a different work.
The seven author-candidate warrants are narrower than many raw answer sentences. They preserve
benchmark, time, population, metric, comparison, edition, retrieval, and judge boundaries. They
are not independently confirmed, and no raw run or profile supplies independence. Four additional
normalized warrant groups remain pending because their captured spans do not close the full
canonical proposition.
Candidate dossier
EM-0029 deterministically projects the accepted packet into
Its content address is
em:dossier:sha256:cbd7a14096a956f642f5c76046d3b49ed648fbe6bf24144c992404a01415af82.
That ID binds candidate bytes; it is not an acceptance or publication receipt.
The candidate keeps distinct:
- eight captured reports from one shared capture program;
- forty-eight citation occurrences, thirty cited URL strings, and twenty-seven resolving URL
- eleven source works, fourteen examined editions, and seventy-two matched exact-span roots;
- seven scoped warrant candidates and zero independently confirmed warrant roots; and
- four pending warrant groups, nine independently rejected claim occurrences, twenty unsupported
roots;
or force-raised occurrences, thirty-four unresolved citations, and three inaccessible carriers.
The encyclopedia reading says the bounded record contains real empirical methods for testing URL
resolution and claim-to-source support, while report agreement adds no independent warrant beyond
the inspected lineages. The skeptical reading says this pilot cannot estimate present-day agent
reliability because the sample is not representative and the unresolved/no-credit set is material.
Both are policy-relative evaluations of the same source graph, not competing source packets.
The first exact-head review found two true but incompletely exposed qualifications. The forward
candidate therefore adds three quote-minimal review-supplement spans: DeepTRACE's 50.3% table
cell, its conflicting 40.3% prose, and the URL-health paper's statement that DRBench supplied
pre-collected model outputs. These spans close the affected sentences but remain separately typed;
they do not turn the accepted 72 EM-0026 exact-span roots into 75 evidence roots.
Material corrections and negative results
Fresh authoritative readback found five important interpretation boundaries:
- link resolution, topical relevance, cited-claim support, and citation coverage are different
- DeepResearch Bench arXiv v1 and ICLR 2026 have materially different FACT values;
- DeepTRACE's own table and prose disagree on the Gemini Deep Research value, 50.3% versus 40.3%;
- *Cited but Not Verified* reuses DeepResearch Bench and BrowseComp query roots, while the
- a resolved claim-to-citation edge is no longer asserted when the target citation remains
measurements and are not pooled;
and
URL-health study reuses DeepResearch Bench outputs; and
unresolved.
The packet also corrects a LiveResearchBench license statement to CC BY-NC-SA 4.0 and records that
Mendeley landing pages were accessible while the credential-free file API was not. These
corrections live in review records; raw answers remain byte-identical.
The machine ledger and reproduction command are documented in
research/how-we-know/agent-citation-lineage/README.md.
Public boundary
The packet may retain the frozen prompt, final answers, citations, and declared public retrieval
receipts. It excludes hidden reasoning, private context, credentials, logged-in state, personal
data, and provider-restricted instructions. Restricted sources receive quote-minimal treatment.
Authority boundary
EM-0026 authorized the accepted source packet. EM-0029 authorizes construction and independent
review of the reversible candidate dossier. Neither task admits a dossier, alters a public lens,
features Case 002, deploys a site, or establishes a universal result about agents. A fresh-clone,
independently rooted reviewer must bind the exact candidate bytes; reproduce source, edition,
span, count, and dependence closure; and verify that neither policy reading strengthens the
accepted evidence.
Public library admission remains a separate EM-0030 gate after EM-0029 review.
The full protocol lives in
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0