Repository object · research-note

GPT-4 bar-exam percentile: preliminary source readiness

Status: HOLD — lead-level mapping only; not a dossier, verdict, admission, or publication recommendation

Source path
research/how-we-know/library-preliminary/gpt-4-bar-exam-percentile.md
Media type
text/markdown
Object ID
em:research-note:sha256:580846ef7efbc1361f7ed29695b7c10565299bd889bcf3cb4d25d4166d162846
Content digest
9fef7d4ce88cac9ff0409255bf780d74d931cbd92faac8cc3d1a3b44fda17119

Source content

GPT-4 bar-exam percentile: preliminary source readiness

Status: **HOLD — lead-level mapping only; not a dossier, verdict, admission, or publication

recommendation**

Task: EM-0027

Source-access window represented: 2026-08-23 UTC

Accepted base for this synthesis: e9ad62b18f21594258643694c55709f78b4f9a50

This packet preserves facts read by a context-isolated source scout from public,

credential-free primary or authoritative sources. Prior model-generated strategy notes receive

no evidentiary credit. “Independent” below means a distinct data or analytical root, not merely

another URL, edition, or author list.

Provisional disposition

HOLD. A later case could examine how an unspecified comparison class became “top 10%,” but

the literal at-launch percentile cannot yet be reconstructed from a disclosed national UBE

total-score distribution. The missing provenance is potentially the case's subject; it is not

permission to fill the gap.

Minimum closure gates are: identify or explicitly represent as unresolved the at-launch

percentile chart and interpolation; freeze the relevant editions and lawful quotation plan;

independently review every dependence edge; and decide whether a later contract is about the

literal percentile claim or comparison-population ambiguity.

Circulated proposition versus reported propositions

| Layer | Source-supported proposition | Boundary |
| --- | --- | --- |
| OpenAI report | A specified GPT-4 evaluation produced a simulated UBE composite of `298/400`, labeled `~90th` percentile and around the top 10% of test takers. | “Test takers” is not “lawyers.” No UBE distribution, jurisdiction, administration, or first-time/repeater composition is cited for the percentile. |
| Katz et al. | A preliminary GPT-4 was evaluated on an MBE practice set and July 2022 MEE/MPT material; the authors report approximately `297` and passage above then-current jurisdiction thresholds. | This was not a live, timed, secure bar administration. OpenAI and Katz report the same collaborator/evaluation root. |
| Martínez | Alternative comparison groups yield materially different modeled ranks; the MBE result can be reproduced more closely than the author-graded essay result. | These are assumption-dependent counter-estimates, not an observed national UBE percentile table or a new full UBE administration. |
| Stronger circulation | Variants such as “beat 90% of lawyers.” | A proposition under review, not a factual conclusion credited by this packet. |

Verified source and access register

Original model report

1. OpenAI, “GPT-4 Technical Report.”

- DOI/identifier: 10.48550/arXiv.2303.08774; arXiv 2303.08774.

- Editions: v1 PDF, submitted 15 March

2023; v6 HTML and

v6 PDF, submitted 4 March 2024;

official unversioned mirror.

- Locators: §3.1/Table 1 and Appendices A.1, A.5, A.7.

- Media/access: HTML and PDF; credential-free.

- Reported anchor: 298/400 (~90th).

- License/treatment: arXiv's distribution license is not a Creative Commons reuse

license; no separate open license was established for the mirror. Link, paraphrase, and

quote minimally; do not redistribute the PDF.

- Lineage: claim/report manifestation of the Katz collaborator experiment, not an

independent performance root.

Underlying score study and artifacts

2. **Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo,

“GPT-4 passes the Bar Exam.”**

- Version of record DOI: 10.1098/rsta.2023.0254; PMCID PMC10894685;

PMID 38403056.

- Access: DOI,

PMC full text, and

Europe PMC XML.

- Edition: received 22 September 2023, accepted 20 December 2023, 2024 version of

record. It contains later discussion and must not be silently treated as the March 2023

launch edition.

- Media/access: HTML/XML credential-free; publisher HTML was automation-hostile.

- License/treatment: CC BY 4.0 for the version of record/PMC.

- Reported anchor: approximately 297; no public July 2022 national percentile table;

a February 2018 Illinois proxy approaches the 90th percentile; plausible comparisons

span roughly 68th–90th depending on timing/state.

- Preprint identity: SSRN DOI 10.2139/ssrn.4389233,

landing page;

automation-sensitive, no open license established, and quote-minimal only.

3. Authors' artifact repository.

- Pinned tree:

90997f740c7197f3f300b013e4345e2ad5621f96;

recursive inventory.

- Inventory: 78 blobs at the inspected tree.

- Missing input: the purchased December 2021 NCBE MBE Complete Practice Exam is not

redistributed; an exact rerun also requires unavailable model/API access.

- License/treatment: no LICENSE at the pinned tree; public access is not reuse

permission. Link-only and quote-minimal.

4. Royal Society Figshare supplement.

- Collection DOI 10.6084/m9.figshare.c.7031287.v1;

article DOI;

article API.

- Inventory: one PDF, rsta20230254_si_001.pdf, 178,633 bytes, MD5

71f8e1e205fb05f847f5a894cc14cf40.

- Media/access: JSON/PDF, credential-free.

- License/treatment: article/file metadata says CC BY 4.0; collection-level license was

null, so do not generalize the file license to the collection.

Authoritative comparison-population sources

5. Illinois Board of Admissions percentile-equivalent charts.

- February 2018:

one-page PDF; 300 -> 90th, 290 -> 85th.

- July 2018:

one-page PDF; 300 -> 70th, 290 -> 59th.

- February 2019:

one-page PDF; 300 -> 90th, 290 -> 83rd.

- Media/access: public PDF, credential-free.

- License/treatment: no open license established; reproduce only necessary numerical

anchors with attribution, not whole charts.

- Gap: neither 297 nor 298 is a printed row, and no interpolation rule or underlying

microdata was disclosed. Katz later names February 2018; Martínez inspected/cited

February 2019. The launch chart remains unidentified.

6. National Conference of Bar Examiners.

- UBE score mechanics: official HTML;

total scale 400, with 50% MBE, 30% MEE, and 20% MPT.

- 2022 MBE distribution:

official HTML; February N=16,504, July N=44,705, total N=61,209.

- 2022 statistics snapshot:

official HTML; estimated February mix 68% repeaters/32% first-timers and July mix

23%/77%, subject to NCBE's classification limits.

- First-time/repeater table:

jurisdiction-reported status, not a national UBE total-score percentile table.

- Media/access: HTML, credential-free.

- License/treatment: no open license established; link and use only necessary attributed

anchors.

Re-analysis

7. Eric Martínez, “Re-evaluating GPT-4's bar exam performance.”

- Version of record DOI 10.1007/s10506-024-09396-9;

full text;

Texas A&M record.

- Edition: accepted 30 January 2024; published 30 March 2024; *Artificial Intelligence

and Law* 33, 581–604.

- Media/access: HTML/PDF, credential-free.

- License/treatment: CC BY 4.0.

- Reported anchors: July comparison about 68th and modeled approximately 62nd among

first-time takers. The version of record is internally inconsistent for the comparison

among passers: Results §3.2.2 reports approximately 45th percentile, while the abstract

and discussion report approximately 48th. Preserve 45th/48th as an unresolved

within-edition discrepancy; both estimates depend on stated normality and

inferred-distribution assumptions.

- Preprint: SSRN DOI 10.2139/ssrn.4441311; no separate open license established.

8. Martínez OSF analysis/code deposit.

- Exact anonymous capability URL:

OSF c8ygu;

API form.

- Inventory: 10 files, 34,906,996 bytes.

- Access: project metadata says public=false, but the article-supplied view-only URL was

anonymously readable. It is an unregistered, capability-dependent deposit.

- License/treatment: no project/file license found; link-only, no redistribution.

- Reproducibility boundary: a fresh run still requires omitted copyrighted questions and

model/API access.

Lineage and counterevidence

  • OpenAI, Katz preprint/VOR, GitHub, and Figshare are manifestations of one experiment root,
  • not four replications.

  • Martínez is a distinct analytical root but reuses the reported score and authoritative
  • comparison aggregates; it is not a new full UBE run.

  • Illinois and NCBE are authoritative comparison-data roots, not model-performance roots.
  • Changing February to July moves the coarse Illinois comparison near the same score by about
  • twenty percentile points.

  • Preserve 298 (OpenAI) and approximately 297 (Katz) as an edition/source discrepancy.
  • Preserve Martínez's approximately 45th versus 48th percentile among-passers statements
  • as an internal version-of-record discrepancy rather than choosing one silently.

  • Preserve distinctions among passing a jurisdictional threshold, percentile rank, essay/MBE
  • component performance, and practicing-lawyer competence.

  • The proprietary snapshot, copyrighted questions, nonofficial essay grading, simulation
  • conditions, prompt protocol, and comparison population all bound generalization.

Negative searches and unresolved artifacts

1. No credential-free authoritative national total-UBE percentile table for the relevant

administration was located in the bounded scout search.

2. OpenAI does not cite a UBE-specific distribution or comparison population.

3. The exact launch chart and interpolation are unresolved.

4. Proprietary model snapshots and purchased MBE questions prevent a full deterministic rerun.

5. Rights remain unresolved for the unlicensed GitHub and OSF artifacts.

6. The OSF deposit depends on a view-only capability URL.

7. SSRN access is automation-sensitive.

8. Official blind essay grading/scaling was not reproduced.

These are bounded-search results, not claims that the artifacts do not exist.

Bounded corpus and cost

  • Core review set: 15 source objects across OpenAI, Katz, Illinois, NCBE, and Martínez.
  • Mechanical inventory: 78 pinned Git blobs, one Figshare PDF, and 10 OSF files
  • (89 files total), not independent evidence.

  • Estimated later cost: 1–2 reviewer-days for core claim/edition/lineage review; 3–5
  • reviewer-days for file-level artifact/method audit.

  • A full rerun is outside scope and currently infeasible without purchased inputs,
  • credentials, and a proprietary snapshot.

The recommendation remains HOLD until the source and comparison-class gaps are explicitly

closed or accepted as represented unknowns under a later research contract.

Build receipt

Reproduce this projection

Reproducible projection
Catalog
em:catalog:sha256:9bfc972213cba2cde167386103dc2c011ee74639fb7f0794c54120fbbdef1a5d
Frontier
em:frontier:sha256:f33be3eae4c75232d56750ef9a1aa79d96274ece3417d65a75c1391bf61a81bf
Accepted commit
f92846570180dfa4511263f8ba98ecd18f7772c9
Epistemic policy
commons-balanced-v0.1
Disclosure policy
public-noninterference-v0.1
Compiler
epistemedia/0.2.0