Repository object · documentation

Case 002 research candidate: agent citation lineage

Status: v1 transport preflight failed and is retained; the v2 packet passed EM-0026 source review; an EM-0029 candidate dossier now exists and awaits exact-head independent review; research only.

Source path
docs/research/case-002-agent-citation-lineage.md
Media type
text/markdown
Object ID
em:documentation:sha256:a16ca25bcf6ad3470b4f122f22424ee40683cf22c0c53f3f5b55510a89f15f42
Content digest
dc444f08c8799ebba4287ac75f7d71ed8697f79a6da7fc72d07e4736aae8fadc

Source content

Case 002 research candidate: agent citation lineage

Status: v1 transport preflight failed and is retained; the v2 packet passed EM-0026 source review;

an EM-0029 candidate dossier now exists and awaits exact-head independent review; research only.

Case 002 asks whether apparent agreement among research agents survives inspection of the sources

and exact spans behind their citations. The pilot is designed to teach a distinct lesson from Case

001: a report, URL, citation, source work, source span, and warranted proposition are different

units.

Frozen target

What empirical evidence published or publicly posted by 2026-08-22 measures whether citations
produced by deep-research agents resolve and actually support the claims made from them?

The dated cutoff prevents a fast-moving product comparison from masquerading as a timeless result.

The question is about public empirical evaluations, not private vendor performance or all agents.

Collection design

The active v2 protocol assigns eight context-isolated runs: four on gpt-5.6-sol and four on

gpt-5.6-terra, each at requested high reasoning effort. Every run receives the same frozen prompt

with no inherited conversation turns. Unavailable provider system, sampling, or retrieval settings

remain unknown.

Before v2 was frozen, the v1 transport preflight detected that one invocation received a prompt

with one extra character. Both started invocations were stopped before final-answer capture. The

failed v1 matrix is not being completed or repaired in place: its exact prompt identities, terminal

statuses, and zero admitted answers remain a negative preflight record. V2 is a separately

identified matrix, not a replacement run inside v1.

Different runtime profiles are not different epistemic observers. The matrix gives the project a

bounded set of outputs to inspect; it does not turn eight reports into eight evidence roots.

What was inspected

For each citation, later work must resolve:

  • requested and final URL, redirect chain, retrieval status, media type, bytes, and digest where
  • an artifact can be captured;

  • source work, edition, exact span, and license treatment;
  • the proposition asserted by the agent and the narrower proposition the span actually supports;
  • modality, causality, scope, time, population, metric, comparison-class, and numerical-strength
  • changes; and

  • shared model, prompt, retrieval, URL, source, span, upstream citation, data, method, and
  • derivation dependencies.

The deterministic research-only ledger currently reports:

8 agent reports → 30 cited URL strings → 27 resolving URL roots → 11 source works → 14
examined editions → 72 matched exact-span roots → 7 author-candidate warrant roots → 0
independently confirmed warrant roots

Every number is computed from captured records and relations. The raw matrix contains 48 citation

occurrences, 127 source-span occurrences, and 52 result claims. Twenty claim occurrences are

unsupported or force-raised after review, including nine whose linked quote fragments do not

support their complete proposition. Thirty-four citation occurrences remain unresolved and

receive no automatic credit. Three cited carriers were inaccessible during

fresh readback: two returned HTTP 403 and PubMed returned a cookie interstitial under HTTP 203. No

URL was malformed or found to identify a different work.

The seven author-candidate warrants are narrower than many raw answer sentences. They preserve

benchmark, time, population, metric, comparison, edition, retrieval, and judge boundaries. They

are not independently confirmed, and no raw run or profile supplies independence. Four additional

normalized warrant groups remain pending because their captured spans do not close the full

canonical proposition.

Candidate dossier

EM-0029 deterministically projects the accepted packet into

candidate-dossier.json.

Its content address is

em:dossier:sha256:cbd7a14096a956f642f5c76046d3b49ed648fbe6bf24144c992404a01415af82.

That ID binds candidate bytes; it is not an acceptance or publication receipt.

The candidate keeps distinct:

  • eight captured reports from one shared capture program;
  • forty-eight citation occurrences, thirty cited URL strings, and twenty-seven resolving URL
  • roots;

  • eleven source works, fourteen examined editions, and seventy-two matched exact-span roots;
  • seven scoped warrant candidates and zero independently confirmed warrant roots; and
  • four pending warrant groups, nine independently rejected claim occurrences, twenty unsupported
  • or force-raised occurrences, thirty-four unresolved citations, and three inaccessible carriers.

The encyclopedia reading says the bounded record contains real empirical methods for testing URL

resolution and claim-to-source support, while report agreement adds no independent warrant beyond

the inspected lineages. The skeptical reading says this pilot cannot estimate present-day agent

reliability because the sample is not representative and the unresolved/no-credit set is material.

Both are policy-relative evaluations of the same source graph, not competing source packets.

The first exact-head review found two true but incompletely exposed qualifications. The forward

candidate therefore adds three quote-minimal review-supplement spans: DeepTRACE's 50.3% table

cell, its conflicting 40.3% prose, and the URL-health paper's statement that DRBench supplied

pre-collected model outputs. These spans close the affected sentences but remain separately typed;

they do not turn the accepted 72 EM-0026 exact-span roots into 75 evidence roots.

Material corrections and negative results

Fresh authoritative readback found five important interpretation boundaries:

  • link resolution, topical relevance, cited-claim support, and citation coverage are different
  • measurements and are not pooled;

  • DeepResearch Bench arXiv v1 and ICLR 2026 have materially different FACT values;
  • DeepTRACE's own table and prose disagree on the Gemini Deep Research value, 50.3% versus 40.3%;
  • and

  • *Cited but Not Verified* reuses DeepResearch Bench and BrowseComp query roots, while the
  • URL-health study reuses DeepResearch Bench outputs; and

  • a resolved claim-to-citation edge is no longer asserted when the target citation remains
  • unresolved.

The packet also corrects a LiveResearchBench license statement to CC BY-NC-SA 4.0 and records that

Mendeley landing pages were accessible while the credential-free file API was not. These

corrections live in review records; raw answers remain byte-identical.

The machine ledger and reproduction command are documented in

research/how-we-know/agent-citation-lineage/README.md.

Public boundary

The packet may retain the frozen prompt, final answers, citations, and declared public retrieval

receipts. It excludes hidden reasoning, private context, credentials, logged-in state, personal

data, and provider-restricted instructions. Restricted sources receive quote-minimal treatment.

Authority boundary

EM-0026 authorized the accepted source packet. EM-0029 authorizes construction and independent

review of the reversible candidate dossier. Neither task admits a dossier, alters a public lens,

features Case 002, deploys a site, or establishes a universal result about agents. A fresh-clone,

independently rooted reviewer must bind the exact candidate bytes; reproduce source, edition,

span, count, and dependence closure; and verify that neither policy reading strengthens the

accepted evidence.

Public library admission remains a separate EM-0030 gate after EM-0029 review.

The full protocol lives in

research/how-we-know/agent-citation-lineage/.

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0