Repository object · research-note
GPT-4 bar-exam percentile research packet
Status: corrected author packet complete; fresh-clone independent re-review pending. This is research input, not an admitted How We Know dossier, public verdict, current-model benchmark, or deployment.
- Media type
text/markdown- Object ID
em:research-note:sha256:202400413ee1cf45826b7f3178a07cdf8203c4e482b92315c07b60b84fd79545- Content digest
26df42f08d79a89e9960bae7b03a11f9cfeab0894cf4f5f57e9022477b571c28
Also filed under
Source content
GPT-4 bar-exam percentile research packet
Status: corrected author packet complete; fresh-clone independent re-review pending. This is
research input, not an admitted How We Know dossier, public verdict, current-model benchmark, or
deployment.
Task: EM-0032
Target question:
How did a historical simulated UBE score reported for GPT-4 become a roughly 90th-percentile
claim, and how does the rank change when the comparison population changes?
Evidence cutoff: 2026-08-27
Current author recommendation
GO, pending independent review, for a later dossier about comparison-class ambiguity and
missing launch provenance. The useful result is not “the real percentile was X.” It is:
- OpenAI's launch report displayed
298/400 (~90th)for “test takers” without identifying the - the later study version of record reports approximately
297, explains how a best-performing - official Illinois February and July charts put the same score region in materially different
- Martínez's assumption-bound re-analysis yields about
62ndamong modeled first-time takers - none of these comparisons ranks GPT-4 against practicing lawyers or establishes current model
UBE chart, administration, denominator, or interpolation;
MBE choice could yield 298 or higher, and treats 68th–90th as a plausible
administration-dependent range;
places;
and about 45th under its encoded passers calculation, while its own abstract and discussion
say about 48th; and
performance or general legal competence.
The original launch chart and interpolation remain unresolved-no-credit. A linear
interpolation is included only as a visible sensitivity calculation; it is never attributed to
OpenAI.
Packet identities
- candidate packet:
- artifact inventory:
- pinned Git body-search manifest:
- 15 preliminary core source objects plus 4 derivation supplements;
- 35 quote-minimal parent spans decomposed into 76 typed cells, clauses, code lines, or contiguous
- 89 mechanical artifacts: 78 pinned Git blobs, 1 Figshare PDF, and 10 OSF files;
- 10 deterministic calculations;
- 5 lineage roots; and
- 10 typed, evidence-bound dependence edges covering data, model, author-social, method, material,
em:research-packet:sha256:535d07e59563b12f66e590c31b0d53a21db1a8dfce1487129a54c5e86b9fd55b;
em:artifact-inventory:sha256:17f52a5509fded7e75b08201f61122d09c7626443857c94f8727c85c0824e61c;
em:git-blob-search:sha256:545908f30ba849c42c860185f92612f4d52a53b4f13b3c3ca5672213e23ba996;
text units;
benchmark, score, comparison-class, citation, and derivation dependence.
All 89 artifacts receive zero automatic independent-evidence credit. OpenAI report editions,
Katz's preprint/VOR/repository/supplement, and the study outputs collapse to one historical model
performance root. Martínez's VOR and OSF deposit collapse to one re-analysis root. The repaired
packet also distinguishes the Katz repository commit from its Git tree, uses the canonical Spring
2022 NCBE testing-column work and visible Rosemary Reshetar byline while retaining conflicting
JSON-LD attribution to Jim Leach, binds every July MBE bin used by the passers calculation, and
keeps the Martínez 45th/48th edition-internal discrepancy visible. Its seven Martínez spans
are bound to an exact Texas A&M institutional PDF capture; current automated refreshes can return
HTTP 403, which remains an explicit carrier limitation rather than being silently substituted.
Files
source-records.json— source, edition, capture, license, span, claim, lineage, negative-search,artifact-inventory.json— content-addressed 89-file metadata inventory;git-blob-search-manifest.json— all 78 pinned Git bodies with SHA-256, UTF-8 search results,candidate-packet.json— deterministic content-addressed packet;build_packet.py— capture helper and offline deterministic builder;normalize_html_visible_text.py— exact-root visible-text normalizer for mutable HTML carriers;verify_git_blob_search.py— pinned commit/tree body-readback and negative-search verifier;verify_packet.py— source/count/math/lineage verifier, adversarial receipt self-test, andindependent-review-receipt.json— absent until a separate reviewer completes exact-head
and limitation records;
and explicit binary no-text-search records;
fail-closed exact-head review gate; and
review.
Source bodies remain outside Git. CC BY works are still quoted minimally; unlicensed works and
artifacts are link/metadata/quote-minimal only.
Reproduce
python research/how-we-know/gpt-4-bar-exam-percentile/build_packet.py --check
python research/how-we-know/gpt-4-bar-exam-percentile/verify_git_blob_search.py \
--repository /path/to/pinned-katz-checkout --check --self-test
python research/how-we-know/gpt-4-bar-exam-percentile/verify_packet.py
python research/how-we-know/gpt-4-bar-exam-percentile/verify_packet.py \
--captures-dir /path/to/exact-html-captures --require-captures --require-review
make check
Before independent review, the third command must fail with independent review receipt missing.
After review, it must bind the exact base, author head and tree, packet bytes, every source,
parent span, typed span unit, calculation, lineage root and edge, command record, clean-state
observation, limitation, and recommendation. A receipt-only child must also bind its Git parent and
tree rather than trusting self-asserted hashes. The review gate also requires fresh raw-to-semantic
recomputation for all five mutable HTML carriers; ordinary offline packet validation does not claim
that an external capture was repeated.
Hard boundary
This packet does not authorize a Case 003 dossier, catalog admission, feature, Pages deployment,
provider call, proprietary rerun, credential use, spend, public launch, or claim about a current
OpenAI model. Those remain separate governed tasks.
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0