Repository object · documentation

How We Know case-selection record

Status: active editorial queue. Originally accepted under EM-0025 and updated under EM-0031 after Cases 001 and 002 became independently reviewed live dossiers.

Source path
docs/editorial/how-we-know-case-roadmap.md
Media type
text/markdown
Object ID
em:documentation:sha256:991bca6f4e19f50ded1b0d782fe7b3fbc5e036e295b80dc07dabef0457df443d
Content digest
c20de70e8a3b417a203284dd48ecff3015fdf3149ed09e73e5888657b1f86bec

Source content

How We Know case-selection record

Status: active editorial queue. Originally accepted under EM-0025 and updated under EM-0031 after

Cases 001 and 002 became independently reviewed live dossiers.

This record synthesizes owner direction and three supplied model-generated strategy and

pre-research packets from 2026-08-22. Those packets are not source artifacts, accepted evidence,

or independent review. Their factual leads are preserved below as retrieval questions, not copied

as public claims or numerical findings.

Decision

Case 002 asks: “When several research agents agree, is that independent evidence?”

Its working title is:

When agent agreement is really one retrieval lineage

The second case must teach a second sentence:

Those were not independent confirmations; they were overlapping retrieval and warrant lineages.

This is the strongest choice because it moves from human publication lineage to agent-output

lineage, speaks directly to the founding product mission, recruits builders and agent researchers,

and can produce one narrow inspectable card without entering politics, health advice, or a mutable

vendor leaderboard.

Case 002 is now live from a reproducible, disclosure-safe trace packet. Its accepted graph retains

unresolved citations, no-credit claims, inaccessible carriers, shared prompt and runtime lineage,

and zero independently confirmed warrant roots rather than converting eight reports into eight

independent observers.

Why the recommendations differed

The supplied editorial perspectives optimized for different goals:

  • The verdict-diversity and distribution argument favored the GPT-4 bar-exam percentile claim. It
  • offers a strong “score reproduced, comparison class misleading” shape and a natural benchmark

    claim template.

  • The product-mission argument favored agent citation laundering. Case 001 already shows that paper
  • titles are not evidence roots; Case 002 should show that agent outputs are not observers merely

    because they arrived through separate runs.

The owner preferred the second direction. The product brief also makes agent-native provenance and

lineage a defining distinction, so agent citation lineage wins Case 002. The bar-exam candidate is

retained at the front of the later library rather than discarded.

Selection rubric

A case becomes stronger when it:

1. teaches a new memorable invariant rather than repeating Case 001's ending;

2. proves the dossier format in a different domain or medium;

3. addresses a reader group the existing case does not already recruit;

4. yields one truthful, screenshotable evidence structure;

5. has a bounded and licensable primary-source corpus;

6. avoids political, medical, personal, or fast-moving vendor claims that become the story;

7. varies the library's verdict shape; and

8. exercises machinery that has not yet been demonstrated.

Importance alone is not a selection criterion.

Candidate ledger

| Candidate | New product lesson | Disposition |
| --- | --- | --- |
| Agent citation lineage | Agent reports, URLs, and citations are not independent observers without retrieval and warrant lineage | **Shipped as Case 002** |
| GPT-4 “90th-percentile bar exam” | A reproduced score can become misleading through an undocumented or wrong comparison class | **Selected for Case 003 research**; no dossier or verdict yet |
| Mehrabian 7–38–55 | The proposition as circulated can be much broader than the proposition tested | **Selected for Case 004 research**; no dossier or verdict yet |
| Do fact-checks work? | Policy views can diverge over outcome, durability, and action while sharing one evidence file | Priority lens-divergence and partial-survivor case |
| Shared dataset, many papers | One dataset may wear many publication identities | Strong reserve; too close to Case 001 as the second proof |
| Growth mindset interventions | Heterogeneity, methods, and incentives can produce dueling meta-analytic readings | Hold until policy and conflict-of-interest handling mature |
| SWE-bench-style headline | A score can depend on contamination, retrieval channels, task validity, and harness | Later eval realm; too mutable and lab-shaped for Case 002 |
| Wikipedia-famous sentence | A page can collapse dissent and provenance | Packaging or comparison surface, not a case by itself |
| Ego depletion | Registered replication and shared-program dependence | Bench: large corpus and familiar story |
| Learning styles | Belief prevalence can remain high after evidentiary failure | Bench: useful later general-audience case |
| Living medical rule | Action-guiding claims need stronger evidence and expertise gates | Not an early case |
| “What is episteme?” | Conceptual orientation, not an empirical second law | Preserved as vocabulary, not a case |

Case 002 research brief

Unit of analysis

The empirical unit is a captured agent trace, not a paper about agent behavior. The later packet

must bind, where available:

  • frozen target question, prompt, system context, model and version, tool configuration, and time;
  • raw agent answer and citation-bearing trace;
  • each requested and resolved URL, redirect chain, retrieval status, and captured edition;
  • each exact cited span and the proposition it actually warrants;
  • claim wording before and after any unsupported increase in modality, scope, causality, time, or
  • numerical precision;

  • shared model, prompt, retrieval, source, span, pretraining, and derivation dependencies; and
  • failed, hallucinated, inaccessible, or unresolved citations.

The target question must be narrow, public, low-risk, and independently answerable from a bounded

source set. A later research contract must choose it before traces are run.

Provisional card grammar

No count is asserted in advance. The candidate card shape is:

N agent reports → U resolving URL roots → S exact span roots → D independent warrant roots

The expanded ledger separately reports hallucinated or unresolved URLs and citations that support

a weaker proposition than the agent's wording. Every number must derive from accepted relations;

none may be maintained as editorial copy.

Policy behavior to test
  • Encyclopedia: describe the answer set, source overlap, and exact supported proposition without
  • calling repeated retrieval independent confirmation.

  • Skeptical: withhold action or strong reliance when agreement collapses to one source, one span,
  • one derivation, unresolved URLs, or unsupported claim strengthening.

The policies should be promoted as divergent only if the compiled outputs materially differ. A

styling-only switch fails the candidate.

Stop conditions

Stop and record a negative result when:

  • raw traces cannot be captured or disclosed without secrets, personal data, or provider-restricted
  • context;

  • source editions and exact spans cannot be re-retrieved and bound;
  • the selected question changes materially during collection;
  • URL uniqueness cannot be separated from source, span, and derivation uniqueness;
  • the output would universalize from one bounded run to all agents; or
  • independent review cannot reproduce the lineage counts.

Retrieval questions retained from pre-research

These are leads for later contracts, not accepted factual statements:

  • Agent citation lineage: identify and retrieve the reported DRBench, ExpertQA, and FORCEBENCH
  • work; verify definitions and rates for nonexistent URLs, non-resolving citations, repeated UGC

    retrieval, and citations whose source supports a weaker claim than the answer.

  • GPT-4 bar exam: retrieve the original model report, underlying score study, authoritative exam
  • population data, and the Martínez re-analysis; reconstruct every comparison population and the

    provenance of the “top 10%” derivation.

  • Fact-check effectiveness: retrieve the four-country experiment, broad evidence review, the
  • 2023 science-misinformation meta-analysis and 2025 reply, plus durability and behavior-boundary

    studies; map shared programs with Case 001 rather than double-counting them.

  • Mehrabian: retrieve both 1967 studies, reconstruct the combination step behind 7–38–55,
  • verify the author's later disclaimer, and sample current authoritative recirculation of the

    broader claim.

  • Shared-dataset cases: inspect Many Analysts, fMRI many-team, immigration many-team, and public
  • cohort examples; choose one association and one file rather than reviewing a field.

  • Growth mindset: retrieve the dueling meta-analyses and methods responses, define conflict-of-
  • interest evidence carefully, and avoid converting contested personal attributions into labels.

  • SWE-bench family: retrieve primary benchmark documentation, issue and task validity records,
  • contamination evidence, retrieval-channel audits, and matched harness results before quoting any

    score delta.

Library roadmap

The goal is a small library before broad publicity, not a single permanent homepage exhibit.

| Slot | Working case | Genre job | Gate |
| --- | --- | --- | --- |
| 001 | Correction repetition and familiarity backfire | Deflates apparent support through participant-data lineage and retains an unresolved root | Shipped |
| 002 | Agent agreement and citation lineage | Makes agent/output dependence visible in the medium Epistemedia serves | Shipped |
| 003 | GPT-4 bar-exam percentile | Technically grounded score, misleading or unresolved comparison-class derivation | EM-0032 research and independent go/hold/fail review |
| 004 | Mehrabian 7–38–55 | General-audience scope-mismatch and proposition-canonicalization case | EM-0033 research and independent go/hold/fail review |
| 005 | Fact-check effectiveness | Partial survivor and genuine policy-divergence case | Resolve meta-analytic frame, outcomes, and shared lineages |
| 006 | Growth mindset interventions | Dueling-meta, heterogeneity, method, and incentive-dependence case | Mature conflict-of-interest and policy handling |

Cases 002–005 form the first target library. They should be researched and admitted separately,

with one task, branch, source packet, independent review, and PR per case. They need not be featured

in numerical order, and no case should wait for a weaker candidate merely to preserve cadence.

Broad-launch gate

Epistemedia's repository and site are already public and live. “Broad launch” means deliberate

promotion, outreach, and presenting the library as a public series; it does not mean first

deployment or first public access.

Broad launch waits until the library contains at least four independently reviewed dossiers,

including admitted and provider-verified Cases 003 and 004. Each case must expose its complete

scoreboard ledger, source path, review boundary, and materially different policy view without

private guidance. Meeting that technical and editorial gate does not automatically authorize

promotion; launch remains a separate owner decision.

Next authority

EM-0032 and EM-0033 separately authorize source-closure research for Cases 003 and 004. Each must

end in an independently reviewed go, hold, or fail result before any dossier or admission task is

registered. Billable runs, paid sources, new accounts, credentials, provider terms, dossier

admission, deployment, and broad promotion remain separate gates.

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0