# Case 003 research candidate: GPT-4's bar-exam percentile

- Object ID: `em:documentation:sha256:14a6353af1d94e333a6c15f848d01b3453c5f71cad9a0c6cdf98b67c9ee62a9d`
- Kind: `documentation`
- Repository path: [`docs/research/case-003-gpt-4-bar-exam-percentile.md`](https://github.com/yoheinakajima/epistemedia/blob/f92846570180dfa4511263f8ba98ecd18f7772c9/docs/research/case-003-gpt-4-bar-exam-percentile.md)
- Content digest: `b0737d145aef503daae8024965de65c358530a722582e740fd219ddef23bb369`

**Also filed under:** [Disclosure and Public Projection](https://epistemedia.org/topics/disclosure/), [Epistemedia](https://epistemedia.org/topics/epistemedia/), [Epistemic Mesh Protocol](https://epistemedia.org/topics/epistemic-mesh/), [Sovereign Realm Federation](https://epistemedia.org/topics/federation/), [Autonomous Governance](https://epistemedia.org/topics/governance/), [Knowledge Objects](https://epistemedia.org/topics/knowledge-objects/), [Human and Agent Interfaces](https://epistemedia.org/topics/public-interfaces/), [Releases and Reproducibility](https://epistemedia.org/topics/releases/), [Research Program](https://epistemedia.org/topics/research-program/), [Security and Adversarial Robustness](https://epistemedia.org/topics/security/)

## Source content

# Case 003 research candidate: GPT-4's bar-exam percentile

Status: corrected content-addressed author research packet complete; independent re-review pending.
This note does not admit or publish Case 003.

The candidate asks a narrower question than “Did GPT-4 pass the bar?” A historical simulated UBE
score was real enough to inspect, but a percentile is not a property of the score alone. It also
depends on who is in the comparison population, when and where they took the exam, and how values
between published score rows are handled.

The current packet preserves four distinct layers:

1. OpenAI's March 2023 report displayed `298/400 (~90th)` and described the result as top 10% of
   “test takers.”
2. The later Katz study reports approximately `297`, explains how a best MBE choice could yield
   `298 or higher`, and says the matching July 2022 national percentile table was not public.
3. Official Illinois charts show that the same score region ranks very differently in February
   and July administrations.
4. Martínez models alternate populations and reports approximately `62nd` among first-time takers;
   its result/table/code support about `45th` under the encoded passers model while its abstract
   and discussion say about `48th`.

The packet does **not** choose an undocumented “true” launch percentile. The exact launch chart,
denominator, and interpolation remain unresolved and receive no credit. Its author recommendation
is to proceed only with a later dossier whose core lesson is comparison-class dependence and
missing claim lineage—not a gotcha about a fake score.

The packet also bars three common overextensions: “test takers” does not mean practicing lawyers;
a simulated historical exam does not establish general legal competence; and none of the captured
results measures a current model.

Its audit structure is deliberately stricter than a bibliography. Thirty-five source spans are
decomposed into 76 typed exact units, including all 21 July MBE score bins used by the modeled
passers calculation. Five underlying roots are connected by ten evidence-bound dependence types,
so a report, repository, supplement, chart, and re-analysis cannot become five independent results
merely by appearing at five URLs. The source record also distinguishes the Katz Git commit from its
tree and identifies Rosemary Reshetar, EdD as the visible author of the Spring 2022 NCBE testing
column while preserving the page's conflicting Jim Leach JSON-LD as metadata, not silent
authorship. A deterministic manifest reads every pinned Git blob body, and mutable NCBE pages are
checked through exact-root semantic normalization rather than raw-byte stability alone.

Full machine-readable records, source identities, limitations, artifact inventory, deterministic
calculations, and the pending independent-review gate are in
[`research/how-we-know/gpt-4-bar-exam-percentile/`](../../research/how-we-know/gpt-4-bar-exam-percentile/README.md).
