Repository object · documentation
Case 003 research candidate: GPT-4's bar-exam percentile
Status: corrected content-addressed author research packet complete; independent re-review pending. This note does not admit or publish Case 003.
- Media type
text/markdown- Object ID
em:documentation:sha256:14a6353af1d94e333a6c15f848d01b3453c5f71cad9a0c6cdf98b67c9ee62a9d- Content digest
b0737d145aef503daae8024965de65c358530a722582e740fd219ddef23bb369
Source content
Case 003 research candidate: GPT-4's bar-exam percentile
Status: corrected content-addressed author research packet complete; independent re-review pending.
This note does not admit or publish Case 003.
The candidate asks a narrower question than “Did GPT-4 pass the bar?” A historical simulated UBE
score was real enough to inspect, but a percentile is not a property of the score alone. It also
depends on who is in the comparison population, when and where they took the exam, and how values
between published score rows are handled.
The current packet preserves four distinct layers:
1. OpenAI's March 2023 report displayed 298/400 (~90th) and described the result as top 10% of
“test takers.”
2. The later Katz study reports approximately 297, explains how a best MBE choice could yield
298 or higher, and says the matching July 2022 national percentile table was not public.
3. Official Illinois charts show that the same score region ranks very differently in February
and July administrations.
4. Martínez models alternate populations and reports approximately 62nd among first-time takers;
its result/table/code support about 45th under the encoded passers model while its abstract
and discussion say about 48th.
The packet does not choose an undocumented “true” launch percentile. The exact launch chart,
denominator, and interpolation remain unresolved and receive no credit. Its author recommendation
is to proceed only with a later dossier whose core lesson is comparison-class dependence and
missing claim lineage—not a gotcha about a fake score.
The packet also bars three common overextensions: “test takers” does not mean practicing lawyers;
a simulated historical exam does not establish general legal competence; and none of the captured
results measures a current model.
Its audit structure is deliberately stricter than a bibliography. Thirty-five source spans are
decomposed into 76 typed exact units, including all 21 July MBE score bins used by the modeled
passers calculation. Five underlying roots are connected by ten evidence-bound dependence types,
so a report, repository, supplement, chart, and re-analysis cannot become five independent results
merely by appearing at five URLs. The source record also distinguishes the Katz Git commit from its
tree and identifies Rosemary Reshetar, EdD as the visible author of the Spring 2022 NCBE testing
column while preserving the page's conflicting Jim Leach JSON-LD as metadata, not silent
authorship. A deterministic manifest reads every pinned Git blob body, and mutable NCBE pages are
checked through exact-root semantic normalization rather than raw-byte stability alone.
Full machine-readable records, source identities, limitations, artifact inventory, deterministic
calculations, and the pending independent-review gate are in
Build receipt
Reproduce this projection
- Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d- Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf- Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9- Epistemic policy
commons-balanced-v0.1- Disclosure policy
public-noninterference-v0.1- Compiler
epistemedia/0.2.0