Repository object · documentation

EM-0026 execution plan

Status: v1 preflight failed and retained; v2 trace matrix and author-side source review complete; exact-head independent review pending.

Source path
docs/execution-plans/EM-0026.md
Media type
text/markdown
Object ID
em:documentation:sha256:7aea72cd3d35a81d11f9eecb750871c4f22f63a19f5679fabf2e83e95ae1229b
Content digest
25600ed83d97c08b17375c5e67ab4c313bd25e4678d0e484db5b2f009e3851fd

Source content

EM-0026 execution plan

Status: v1 preflight failed and retained; v2 trace matrix and author-side source review complete;

exact-head independent review pending.

Objective

Capture a bounded, disclosure-safe set of agent research answers and determine whether their

apparent citation agreement collapses under URL, source-work, edition, exact-span, and warrant

lineage inspection.

Authority and accepted base

  • immutable task: tasks/contracts/EM-0026.json;
  • accepted base: 82b56e407712818ced9e9a9561eb35bd6695ff5d;
  • selected direction: docs/editorial/how-we-know-case-roadmap.md;
  • allowed changes: the EM-0026 research packet, this plan, the research summary, and runs/**.

Sequence

1. Freeze the target decision, cutoff, exact prompt, eight-slot matrix, trace envelope, disclosure

boundary, dependence dimensions, count grammar, and stop conditions. Complete for v2.

2. Bind those files and the failed v1 preflight identity to one exact commit before generating the

first admissible answer. Satisfied by the v2 freeze commit containing these inputs.

3. Run exactly eight fresh-context captures, retaining failed or incomplete slots without

replacement. Complete: eight terminal v2 records.

4. Normalize citations without changing the raw answer bytes. **Complete: 48 citation occurrences

and 30 distinct requested URLs normalized.**

5. Independently retrieve primary or authoritative source editions and record exact spans,

licensing, and unavailable artifacts. **Complete for author review: 27 usable public URL

readbacks, three inaccessible carriers, 14 editions, 127 inspected span occurrences, 72 matched

span roots, and 34 unresolved citation occurrences.**

6. Compare agent propositions with source spans and model any unsupported increase in modality,

causality, scope, time, population, comparison class, or numerical precision. **Complete as 52

occurrence reviews yielding seven narrower author-candidate warrants, four pending warrant

groups, and seven explicit correction records. Nine over-credited claim occurrences now

receive no warrant credit.**

7. Derive report, URL, work, edition, span, warrant, invalid, and unresolved counts from relations.

Complete and reproducible with build_evidence_ledger.py verify.

8. Obtain a fresh-clone independent review of every trace identity, source readback, span match,

dependence edge, and displayed count.

9. Run full deterministic repository validation and submit the result as research only.

Fixed controls

  • question cutoff: 2026-08-22;
  • runs: exactly eight, four per requested runtime profile;
  • prompt: byte-identical across every run;
  • inherited conversation turns: zero;
  • retrieval: credential-free public reads only;
  • replacement runs: forbidden;
  • independence: never granted from a run, agent label, profile, URL, or citation alone;
  • unknown settings: recorded as unknown, never inferred;
  • publication/admission/deployment: not authorized.

Current checkpoint

The target screen passes with explicit time and vendor-volatility boundaries. The prompt is neutral

about success or failure and requests primary sources, exact locators, quote-minimal spans,

quantitative basis, scope, counterevidence, and unknowns. All eight v2 slots completed with the

frozen prompt and disclosure-safe records. They remain observations sharing prompt and runtime

lineage, not eight independent evidence roots.

The v1 transport preflight exposed a one-character prompt mismatch between two started

invocations. Both were interrupted before final-answer capture; zero answers and zero citations

were admitted. The immutable failure record prevents those slots from being silently retried or

treated as part of a clean matrix. A separately identified v2 prompt, matrix, envelope, and

fail-closed verifier are co-versioned in one freeze commit. Every later trace must name that exact

commit before its answer is admissible. The later review preserves the answer and trace bytes while

adding URL readbacks, work and edition normalization, exact-span match records, proposition-level

warrant reviews, dependence edges, and corrections.

Current relation-derived counts are 8 reports, 48 citation occurrences, 30 cited URL strings, 27

resolving URL roots, 11 source works, 14 examined editions, 127 raw span occurrences, 72 matched

exact-span roots, 52 raw claim occurrences, seven author-candidate warrant roots, 34 unresolved

citation occurrences, 20 unsupported or force-raised claim occurrences, and zero independently

confirmed warrant roots. No empirical finding has

been accepted, and no Case 002 dossier exists.

The author review retains these material limitations rather than editing the runs:

  • two cited carriers returned HTTP 403 and PubMed returned an HTTP-203 cookie interstitial;
  • supplementary Mendeley file bytes were not captured because the credential-free API returned
  • HTTP 401;

  • LiveResearchBench's raw license statement was corrected to CC BY-NC-SA 4.0;
  • one run omitted BrowseComp from *Cited but Not Verified*'s query sources;
  • DeepResearch Bench v1 and ICLR values are edition-specific; and
  • DeepTRACE Table 1 and prose disagree on one Gemini value;
  • nine claim occurrences receive no credit because their linked spans establish only part of the
  • asserted method, comparison, metric, scope, or direction; and

  • four warrant groups remain pending because their captured spans do not semantically close the
  • normalized proposition. Claim-to-citation edges now inherit the target citation's actual

    resolved or unresolved status.

The next acceptance step is a fresh-clone independent exact-head review. That review must retrieve

new public editions, inspect the unresolved set, reproduce every identity and count, and return a

PASS or CHANGES REQUIRED receipt. This author branch cannot set the independent warrant count above

zero or approve itself.

Validation

  • parse every JSON control file;
  • verify v2 prompt bytes and digest, failed-v1 record, and superseded-v1 protocol identities;
  • verify the matrix contains exactly four gpt-5.6-sol and four gpt-5.6-terra slots;
  • verify every trace uses the exact prompt and retains terminal failures;
  • run python research/how-we-know/agent-citation-lineage/build_evidence_ledger.py verify;
  • compare changed paths with EM-0026 authority;
  • run git diff --check and make check PYTHON=.venv/bin/python;
  • require exact-head independent review before protected merge.

Review correction record

The first complete independent semantic pass reproduced all 72 credited span roots, then rejected

nine over-credited claim occurrences, four candidate-warrant groups, and 44 falsely resolved

claim-to-citation edges. The raw eight answers and traces remain byte-identical. The corrected

ledger derives seven candidate warrant roots, 20 unsupported or force-raised claim occurrences,

and explicit unresolved edge states. Author validation receipt:

runs/proposals/20260823T181539Z-9f2fc4bda331.json. A fresh exact-head independent receipt remains

required before merge.

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0