Repository object · documentation
EM-0026 execution plan
Status: v1 preflight failed and retained; v2 trace matrix and author-side source review complete; exact-head independent review pending.
- Source path
docs/execution-plans/EM-0026.md- Media type
text/markdown- Object ID
em:documentation:sha256:7aea72cd3d35a81d11f9eecb750871c4f22f63a19f5679fabf2e83e95ae1229b- Content digest
25600ed83d97c08b17375c5e67ab4c313bd25e4678d0e484db5b2f009e3851fd
Source content
EM-0026 execution plan
Status: v1 preflight failed and retained; v2 trace matrix and author-side source review complete;
exact-head independent review pending.
Objective
Capture a bounded, disclosure-safe set of agent research answers and determine whether their
apparent citation agreement collapses under URL, source-work, edition, exact-span, and warrant
lineage inspection.
Authority and accepted base
- immutable task:
tasks/contracts/EM-0026.json; - accepted base:
82b56e407712818ced9e9a9561eb35bd6695ff5d; - selected direction:
docs/editorial/how-we-know-case-roadmap.md; - allowed changes: the EM-0026 research packet, this plan, the research summary, and
runs/**.
Sequence
1. Freeze the target decision, cutoff, exact prompt, eight-slot matrix, trace envelope, disclosure
boundary, dependence dimensions, count grammar, and stop conditions. Complete for v2.
2. Bind those files and the failed v1 preflight identity to one exact commit before generating the
first admissible answer. Satisfied by the v2 freeze commit containing these inputs.
3. Run exactly eight fresh-context captures, retaining failed or incomplete slots without
replacement. Complete: eight terminal v2 records.
4. Normalize citations without changing the raw answer bytes. **Complete: 48 citation occurrences
and 30 distinct requested URLs normalized.**
5. Independently retrieve primary or authoritative source editions and record exact spans,
licensing, and unavailable artifacts. **Complete for author review: 27 usable public URL
readbacks, three inaccessible carriers, 14 editions, 127 inspected span occurrences, 72 matched
span roots, and 34 unresolved citation occurrences.**
6. Compare agent propositions with source spans and model any unsupported increase in modality,
causality, scope, time, population, comparison class, or numerical precision. **Complete as 52
occurrence reviews yielding seven narrower author-candidate warrants, four pending warrant
groups, and seven explicit correction records. Nine over-credited claim occurrences now
receive no warrant credit.**
7. Derive report, URL, work, edition, span, warrant, invalid, and unresolved counts from relations.
Complete and reproducible with build_evidence_ledger.py verify.
8. Obtain a fresh-clone independent review of every trace identity, source readback, span match,
dependence edge, and displayed count.
9. Run full deterministic repository validation and submit the result as research only.
Fixed controls
- question cutoff: 2026-08-22;
- runs: exactly eight, four per requested runtime profile;
- prompt: byte-identical across every run;
- inherited conversation turns: zero;
- retrieval: credential-free public reads only;
- replacement runs: forbidden;
- independence: never granted from a run, agent label, profile, URL, or citation alone;
- unknown settings: recorded as
unknown, never inferred; - publication/admission/deployment: not authorized.
Current checkpoint
The target screen passes with explicit time and vendor-volatility boundaries. The prompt is neutral
about success or failure and requests primary sources, exact locators, quote-minimal spans,
quantitative basis, scope, counterevidence, and unknowns. All eight v2 slots completed with the
frozen prompt and disclosure-safe records. They remain observations sharing prompt and runtime
lineage, not eight independent evidence roots.
The v1 transport preflight exposed a one-character prompt mismatch between two started
invocations. Both were interrupted before final-answer capture; zero answers and zero citations
were admitted. The immutable failure record prevents those slots from being silently retried or
treated as part of a clean matrix. A separately identified v2 prompt, matrix, envelope, and
fail-closed verifier are co-versioned in one freeze commit. Every later trace must name that exact
commit before its answer is admissible. The later review preserves the answer and trace bytes while
adding URL readbacks, work and edition normalization, exact-span match records, proposition-level
warrant reviews, dependence edges, and corrections.
Current relation-derived counts are 8 reports, 48 citation occurrences, 30 cited URL strings, 27
resolving URL roots, 11 source works, 14 examined editions, 127 raw span occurrences, 72 matched
exact-span roots, 52 raw claim occurrences, seven author-candidate warrant roots, 34 unresolved
citation occurrences, 20 unsupported or force-raised claim occurrences, and zero independently
confirmed warrant roots. No empirical finding has
been accepted, and no Case 002 dossier exists.
The author review retains these material limitations rather than editing the runs:
- two cited carriers returned HTTP 403 and PubMed returned an HTTP-203 cookie interstitial;
- supplementary Mendeley file bytes were not captured because the credential-free API returned
- LiveResearchBench's raw license statement was corrected to CC BY-NC-SA 4.0;
- one run omitted BrowseComp from *Cited but Not Verified*'s query sources;
- DeepResearch Bench v1 and ICLR values are edition-specific; and
- DeepTRACE Table 1 and prose disagree on one Gemini value;
- nine claim occurrences receive no credit because their linked spans establish only part of the
- four warrant groups remain pending because their captured spans do not semantically close the
HTTP 401;
asserted method, comparison, metric, scope, or direction; and
normalized proposition. Claim-to-citation edges now inherit the target citation's actual
resolved or unresolved status.
The next acceptance step is a fresh-clone independent exact-head review. That review must retrieve
new public editions, inspect the unresolved set, reproduce every identity and count, and return a
PASS or CHANGES REQUIRED receipt. This author branch cannot set the independent warrant count above
zero or approve itself.
Validation
- parse every JSON control file;
- verify v2 prompt bytes and digest, failed-v1 record, and superseded-v1 protocol identities;
- verify the matrix contains exactly four
gpt-5.6-soland fourgpt-5.6-terraslots; - verify every trace uses the exact prompt and retains terminal failures;
- run
python research/how-we-know/agent-citation-lineage/build_evidence_ledger.py verify; - compare changed paths with EM-0026 authority;
- run
git diff --checkandmake check PYTHON=.venv/bin/python; - require exact-head independent review before protected merge.
Review correction record
The first complete independent semantic pass reproduced all 72 credited span roots, then rejected
nine over-credited claim occurrences, four candidate-warrant groups, and 44 falsely resolved
claim-to-citation edges. The raw eight answers and traces remain byte-identical. The corrected
ledger derives seven candidate warrant roots, 20 unsupported or force-raised claim occurrences,
and explicit unresolved edge states. Author validation receipt:
runs/proposals/20260823T181539Z-9f2fc4bda331.json. A fresh exact-head independent receipt remains
required before merge.
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0