Repository object · documentation
How We Know case-selection record
Status: active editorial queue. Originally accepted under EM-0025 and updated under EM-0031 after Cases 001 and 002 became independently reviewed live dossiers.
- Source path
docs/editorial/how-we-know-case-roadmap.md- Media type
text/markdown- Object ID
em:documentation:sha256:991bca6f4e19f50ded1b0d782fe7b3fbc5e036e295b80dc07dabef0457df443d- Content digest
c20de70e8a3b417a203284dd48ecff3015fdf3149ed09e73e5888657b1f86bec
Source content
How We Know case-selection record
Status: active editorial queue. Originally accepted under EM-0025 and updated under EM-0031 after
Cases 001 and 002 became independently reviewed live dossiers.
This record synthesizes owner direction and three supplied model-generated strategy and
pre-research packets from 2026-08-22. Those packets are not source artifacts, accepted evidence,
or independent review. Their factual leads are preserved below as retrieval questions, not copied
as public claims or numerical findings.
Decision
Case 002 asks: “When several research agents agree, is that independent evidence?”
Its working title is:
When agent agreement is really one retrieval lineage
The second case must teach a second sentence:
Those were not independent confirmations; they were overlapping retrieval and warrant lineages.
This is the strongest choice because it moves from human publication lineage to agent-output
lineage, speaks directly to the founding product mission, recruits builders and agent researchers,
and can produce one narrow inspectable card without entering politics, health advice, or a mutable
vendor leaderboard.
Case 002 is now live from a reproducible, disclosure-safe trace packet. Its accepted graph retains
unresolved citations, no-credit claims, inaccessible carriers, shared prompt and runtime lineage,
and zero independently confirmed warrant roots rather than converting eight reports into eight
independent observers.
Why the recommendations differed
The supplied editorial perspectives optimized for different goals:
- The verdict-diversity and distribution argument favored the GPT-4 bar-exam percentile claim. It
- The product-mission argument favored agent citation laundering. Case 001 already shows that paper
offers a strong “score reproduced, comparison class misleading” shape and a natural benchmark
claim template.
titles are not evidence roots; Case 002 should show that agent outputs are not observers merely
because they arrived through separate runs.
The owner preferred the second direction. The product brief also makes agent-native provenance and
lineage a defining distinction, so agent citation lineage wins Case 002. The bar-exam candidate is
retained at the front of the later library rather than discarded.
Selection rubric
A case becomes stronger when it:
1. teaches a new memorable invariant rather than repeating Case 001's ending;
2. proves the dossier format in a different domain or medium;
3. addresses a reader group the existing case does not already recruit;
4. yields one truthful, screenshotable evidence structure;
5. has a bounded and licensable primary-source corpus;
6. avoids political, medical, personal, or fast-moving vendor claims that become the story;
7. varies the library's verdict shape; and
8. exercises machinery that has not yet been demonstrated.
Importance alone is not a selection criterion.
Candidate ledger
| Candidate | New product lesson | Disposition |
| --- | --- | --- |
| Agent citation lineage | Agent reports, URLs, and citations are not independent observers without retrieval and warrant lineage | **Shipped as Case 002** |
| GPT-4 “90th-percentile bar exam” | A reproduced score can become misleading through an undocumented or wrong comparison class | **Selected for Case 003 research**; no dossier or verdict yet |
| Mehrabian 7–38–55 | The proposition as circulated can be much broader than the proposition tested | **Selected for Case 004 research**; no dossier or verdict yet |
| Do fact-checks work? | Policy views can diverge over outcome, durability, and action while sharing one evidence file | Priority lens-divergence and partial-survivor case |
| Shared dataset, many papers | One dataset may wear many publication identities | Strong reserve; too close to Case 001 as the second proof |
| Growth mindset interventions | Heterogeneity, methods, and incentives can produce dueling meta-analytic readings | Hold until policy and conflict-of-interest handling mature |
| SWE-bench-style headline | A score can depend on contamination, retrieval channels, task validity, and harness | Later eval realm; too mutable and lab-shaped for Case 002 |
| Wikipedia-famous sentence | A page can collapse dissent and provenance | Packaging or comparison surface, not a case by itself |
| Ego depletion | Registered replication and shared-program dependence | Bench: large corpus and familiar story |
| Learning styles | Belief prevalence can remain high after evidentiary failure | Bench: useful later general-audience case |
| Living medical rule | Action-guiding claims need stronger evidence and expertise gates | Not an early case |
| “What is episteme?” | Conceptual orientation, not an empirical second law | Preserved as vocabulary, not a case |
Case 002 research brief
Unit of analysis
The empirical unit is a captured agent trace, not a paper about agent behavior. The later packet
must bind, where available:
- frozen target question, prompt, system context, model and version, tool configuration, and time;
- raw agent answer and citation-bearing trace;
- each requested and resolved URL, redirect chain, retrieval status, and captured edition;
- each exact cited span and the proposition it actually warrants;
- claim wording before and after any unsupported increase in modality, scope, causality, time, or
- shared model, prompt, retrieval, source, span, pretraining, and derivation dependencies; and
- failed, hallucinated, inaccessible, or unresolved citations.
numerical precision;
The target question must be narrow, public, low-risk, and independently answerable from a bounded
source set. A later research contract must choose it before traces are run.
Provisional card grammar
No count is asserted in advance. The candidate card shape is:
N agent reports → U resolving URL roots → S exact span roots → D independent warrant roots
The expanded ledger separately reports hallucinated or unresolved URLs and citations that support
a weaker proposition than the agent's wording. Every number must derive from accepted relations;
none may be maintained as editorial copy.
Policy behavior to test
- Encyclopedia: describe the answer set, source overlap, and exact supported proposition without
- Skeptical: withhold action or strong reliance when agreement collapses to one source, one span,
calling repeated retrieval independent confirmation.
one derivation, unresolved URLs, or unsupported claim strengthening.
The policies should be promoted as divergent only if the compiled outputs materially differ. A
styling-only switch fails the candidate.
Stop conditions
Stop and record a negative result when:
- raw traces cannot be captured or disclosed without secrets, personal data, or provider-restricted
- source editions and exact spans cannot be re-retrieved and bound;
- the selected question changes materially during collection;
- URL uniqueness cannot be separated from source, span, and derivation uniqueness;
- the output would universalize from one bounded run to all agents; or
- independent review cannot reproduce the lineage counts.
context;
Retrieval questions retained from pre-research
These are leads for later contracts, not accepted factual statements:
- Agent citation lineage: identify and retrieve the reported DRBench, ExpertQA, and FORCEBENCH
- GPT-4 bar exam: retrieve the original model report, underlying score study, authoritative exam
- Fact-check effectiveness: retrieve the four-country experiment, broad evidence review, the
- Mehrabian: retrieve both 1967 studies, reconstruct the combination step behind 7–38–55,
- Shared-dataset cases: inspect Many Analysts, fMRI many-team, immigration many-team, and public
- Growth mindset: retrieve the dueling meta-analyses and methods responses, define conflict-of-
- SWE-bench family: retrieve primary benchmark documentation, issue and task validity records,
work; verify definitions and rates for nonexistent URLs, non-resolving citations, repeated UGC
retrieval, and citations whose source supports a weaker claim than the answer.
population data, and the Martínez re-analysis; reconstruct every comparison population and the
provenance of the “top 10%” derivation.
2023 science-misinformation meta-analysis and 2025 reply, plus durability and behavior-boundary
studies; map shared programs with Case 001 rather than double-counting them.
verify the author's later disclaimer, and sample current authoritative recirculation of the
broader claim.
cohort examples; choose one association and one file rather than reviewing a field.
interest evidence carefully, and avoid converting contested personal attributions into labels.
contamination evidence, retrieval-channel audits, and matched harness results before quoting any
score delta.
Library roadmap
The goal is a small library before broad publicity, not a single permanent homepage exhibit.
| Slot | Working case | Genre job | Gate |
| --- | --- | --- | --- |
| 001 | Correction repetition and familiarity backfire | Deflates apparent support through participant-data lineage and retains an unresolved root | Shipped |
| 002 | Agent agreement and citation lineage | Makes agent/output dependence visible in the medium Epistemedia serves | Shipped |
| 003 | GPT-4 bar-exam percentile | Technically grounded score, misleading or unresolved comparison-class derivation | EM-0032 research and independent go/hold/fail review |
| 004 | Mehrabian 7–38–55 | General-audience scope-mismatch and proposition-canonicalization case | EM-0033 research and independent go/hold/fail review |
| 005 | Fact-check effectiveness | Partial survivor and genuine policy-divergence case | Resolve meta-analytic frame, outcomes, and shared lineages |
| 006 | Growth mindset interventions | Dueling-meta, heterogeneity, method, and incentive-dependence case | Mature conflict-of-interest and policy handling |
Cases 002–005 form the first target library. They should be researched and admitted separately,
with one task, branch, source packet, independent review, and PR per case. They need not be featured
in numerical order, and no case should wait for a weaker candidate merely to preserve cadence.
Broad-launch gate
Epistemedia's repository and site are already public and live. “Broad launch” means deliberate
promotion, outreach, and presenting the library as a public series; it does not mean first
deployment or first public access.
Broad launch waits until the library contains at least four independently reviewed dossiers,
including admitted and provider-verified Cases 003 and 004. Each case must expose its complete
scoreboard ledger, source path, review boundary, and materially different policy view without
private guidance. Meeting that technical and editorial gate does not automatically authorize
promotion; launch remains a separate owner decision.
Next authority
EM-0032 and EM-0033 separately authorize source-closure research for Cases 003 and 004. Each must
end in an independently reviewed go, hold, or fail result before any dossier or admission task is
registered. Billable runs, paid sources, new accounts, credentials, provider terms, dossier
admission, deployment, and broad promotion remain separate gates.
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0