Repository object · documentation

Case 003 research candidate: GPT-4's bar-exam percentile

Status: corrected content-addressed author research packet complete; independent re-review pending. This note does not admit or publish Case 003.

Source path
docs/research/case-003-gpt-4-bar-exam-percentile.md
Media type
text/markdown
Object ID
em:documentation:sha256:14a6353af1d94e333a6c15f848d01b3453c5f71cad9a0c6cdf98b67c9ee62a9d
Content digest
b0737d145aef503daae8024965de65c358530a722582e740fd219ddef23bb369

Source content

Case 003 research candidate: GPT-4's bar-exam percentile

Status: corrected content-addressed author research packet complete; independent re-review pending.

This note does not admit or publish Case 003.

The candidate asks a narrower question than “Did GPT-4 pass the bar?” A historical simulated UBE

score was real enough to inspect, but a percentile is not a property of the score alone. It also

depends on who is in the comparison population, when and where they took the exam, and how values

between published score rows are handled.

The current packet preserves four distinct layers:

1. OpenAI's March 2023 report displayed 298/400 (~90th) and described the result as top 10% of

“test takers.”

2. The later Katz study reports approximately 297, explains how a best MBE choice could yield

298 or higher, and says the matching July 2022 national percentile table was not public.

3. Official Illinois charts show that the same score region ranks very differently in February

and July administrations.

4. Martínez models alternate populations and reports approximately 62nd among first-time takers;

its result/table/code support about 45th under the encoded passers model while its abstract

and discussion say about 48th.

The packet does not choose an undocumented “true” launch percentile. The exact launch chart,

denominator, and interpolation remain unresolved and receive no credit. Its author recommendation

is to proceed only with a later dossier whose core lesson is comparison-class dependence and

missing claim lineage—not a gotcha about a fake score.

The packet also bars three common overextensions: “test takers” does not mean practicing lawyers;

a simulated historical exam does not establish general legal competence; and none of the captured

results measures a current model.

Its audit structure is deliberately stricter than a bibliography. Thirty-five source spans are

decomposed into 76 typed exact units, including all 21 July MBE score bins used by the modeled

passers calculation. Five underlying roots are connected by ten evidence-bound dependence types,

so a report, repository, supplement, chart, and re-analysis cannot become five independent results

merely by appearing at five URLs. The source record also distinguishes the Katz Git commit from its

tree and identifies Rosemary Reshetar, EdD as the visible author of the Spring 2022 NCBE testing

column while preserving the page's conflicting Jim Leach JSON-LD as metadata, not silent

authorship. A deterministic manifest reads every pinned Git blob body, and mutable NCBE pages are

checked through exact-root semantic normalization rather than raw-byte stability alone.

Full machine-readable records, source identities, limitations, artifact inventory, deterministic

calculations, and the pending independent-review gate are in

research/how-we-know/gpt-4-bar-exam-percentile/.

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0